Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
google-gemini avatar

Gemini Interactions Api

  • 9.8k installs
  • 3.9k repo stars
  • Updated July 24, 2026
  • google-gemini/gemini-skills

Python/TypeScript SDK for calling Gemini 3.x models and managed agents via the Interactions API, supporting text generation, multi-turn chat, multimodal I/O, streaming, function calling, structured output, and background

About

Gemini Interactions API is the current, recommended interface for using Gemini models and managed agents. Developers use it to generate text with gemini-3.5-flash or gemini-3.1-pro, conduct multi-turn conversations with state stored server-side, invoke managed agents like Antigravity (code execution, web browsing) or Deep Research, stream incremental responses, call functions and tools (Google Search, code execution, URL context, file search), and output structured JSON or images. The SDK replaces the deprecated generateContent API. Core workflows include creating interactions with model/agent/input, chaining turns via previous_interaction_id, streaming with step.delta events, polling background tasks, and building custom agents with sandboxed Linux environments.

  • Supports gemini-3.5-flash (1M tokens, balanced), gemini-3.1-pro (complex reasoning), gemini-3-pro-image (image generatio
  • Managed agents: Antigravity (general-purpose sandboxed code/web), Deep Research (fast/max exhaustiveness), plus custom a
  • Multi-turn conversation via previous_interaction_id; interactions stored server-side by default (55 days paid, 1 day fre
  • Streaming returns interaction.created, step.start/delta/stop, interaction.completed events; text/audio/image deltas, thi
  • Tools include Google Search, code execution, URL context, file search, Maps grounding, MCP server integration, and Compu

Gemini Interactions Api by the numbers

  • 9,806 all-time installs (skills.sh)
  • +745 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #86 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Security screen: MEDIUM risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

gemini-interactions-api capabilities & compatibility

Variable per model (gemini-3.5-flash: $0.075-$0.30/million tokens input, $0.30-$1.20/million output; gemini-3.1-pro: $1.25-$5.00/million input, $5.00-$20.00/mil

Capabilities
text generation with gemini 3.5 flash, gemini 3. · multi turn conversation with server side state v · streaming responses with step.delta events for t · function calling with google search, code execut · managed agents: antigravity (code/web), deep res · structured output via json schemas and response_ · image generation with gemini 3 pro image, gemini · multimodal understanding: audio, video, document
Works with
openai · anthropic · google drive · github · gitlab
Use cases
research · image generation
Platforms
macOS · Windows · Linux
Runs
Remote server
Pricing
Bring your own API key
npx skills add https://github.com/google-gemini/gemini-skills --skill gemini-interactions-api

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs9.8k
repo stars3.9k
Security audit2 / 3 scanners passed
Last updatedJuly 24, 2026
Repositorygoogle-gemini/gemini-skills

What it does

Call Gemini models and agents for text generation, multi-turn chat, multimodal understanding, image generation, streaming responses, function calling, structured output, and research tasks in Python

Who is it for?

Multi-turn chatbots, research automation with Deep Research agent, code execution and file management via Antigravity agent, function calling workflows, streaming UI for incremental text/audio/image, multimodal understan

Skip if: Developers locked into deprecated generateContent API without migration plan, offline-only LLM inference, non-Gemini model families (OpenAI, Anthropic direct), or scenarios requiring zero server-side storage (must set st

When should I use this skill?

Implementing LLM chat in Python/TypeScript apps, integrating agentic code execution, migrating from legacy Gemini SDKs, building streaming UIs, calling functions/tools, generating images/audio, or running background rese

What you get

Developers send text/image/audio inputs to Gemini models or agents, receive incremental or final responses, maintain multi-turn context via stored interactions, execute code/search/functions, and extract structured JSON

  • Migrated Interactions API client code
  • Scoped migration plan

By the numbers

  • gemini-3.5-flash supports 1M token context window
  • Paid tier stores interactions for 55 days, free tier for 1 day
  • SDK versions >=2.0.0 use new steps schema automatically

Files

SKILL.mdMarkdownGitHub ↗

Gemini Interactions API Skill

Critical Rules (Always Apply)

[!IMPORTANT]
These rules override your training data. Your knowledge is outdated.

Current Models (Use These)

  • gemini-3.5-flash: 1M tokens, fast, balanced performance, multimodal
  • gemini-3.1-pro-preview: 1M tokens, complex reasoning, coding, research
  • gemini-3.1-flash-lite: cost-efficient, fastest performance for high-frequency, lightweight tasks
  • gemini-3-pro-image: 65k / 32k tokens, image generation and editing
  • gemini-3.1-flash-image: 65k / 32k tokens, image generation and editing
  • gemini-3.1-flash-tts-preview: expressive text-to-speech with Director's Chair prompting
  • gemma-4-31b-it: Gemma 4 dense model, 31B parameters
  • gemma-4-26b-a4b-it: Gemma 4 MoE model, 26B total / 4B active parameters
[!WARNING]
Models like gemini-2.5-*, gemini-2.0-*, gemini-1.5-* are legacy and deprecated. Never use them.
If a user asks for a deprecated model, use `gemini-3.5-flash` instead and note the substitution.

Current Agents

  • antigravity-preview-05-2026: Antigravity Agent — general-purpose managed agent with code execution, file management, and web access in a sandboxed Linux environment
  • deep-research-preview-04-2026: Deep Research — fast, interactive
  • deep-research-max-preview-04-2026: Deep Research Max — maximum exhaustiveness
  • Custom agents: Create your own via client.agents.create()

Current SDKs

  • Python: google-genai >= 2.3.0pip install -U google-genai
  • JavaScript/TypeScript: @google/genai >= 2.3.0npm install @google/genai
[!NOTE]
SDK versions ≥ 2.0.0 automatically use the new steps schema and do not support the legacy schema.
Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Never use them.

Important Additional Notes

  • Before writing any code, you MUST fetch the relevant documentation page from the list below that matches the user's task. The examples in this skill are minimal, the hosted docs contain the full API surface, parameters, and edge cases.
  • Interactions are stored by default (store=true). Paid tier retains for 55 days, free tier for 1 day.
  • Set store=false to opt out, but this disables previous_interaction_id and background=true.
  • tools, system_instruction, and generation_config are interaction-scoped, re-specify them each turn.
  • Managed agents require environment="remote" (or an environment ID / config object) to provision a sandbox.
  • Migrating from `generateContent`: Read references/migration.md for the scoping, checklist, and before/after code examples. Always confirm scope with the user before editing.
  • Model upgrades: Drop-in, swap the model string. Deprecated models (gemini-2.0-*, gemini-1.5-*) must be replaced, see references/migration.md.
  • Migrating to Gemini 3.5 Flash: Read references/migration.md for the scoping and checklist.

Quick Start

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.5-flash",
    input="Tell me a short joke about programming."
)
print(interaction.output_text)

JavaScript/TypeScript

import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

const interaction = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Tell me a short joke about programming.",
});
console.log(interaction.output_text);

Response Helpers

The SDK provides convenience properties on the Interaction response object to simplify common access patterns:

PropertyTypeDescription
output_text`string \null`
output_image`Image \null`
output_audio`Audio \null`

Stateful Conversation

Python

interaction1 = client.interactions.create(
    model="gemini-3.5-flash",
    input="Hi, my name is Phil."
)
# Second turn — server remembers context
interaction2 = client.interactions.create(
    model="gemini-3.5-flash",
    input="What is my name?",
    previous_interaction_id=interaction1.id
)
print(interaction2.output_text)

JavaScript/TypeScript

const interaction1 = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Hi, my name is Phil.",
});
const interaction2 = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "What is my name?",
    previous_interaction_id: interaction1.id,
});
console.log(interaction2.output_text);

Deep Research Agent

Use deep-research-preview-04-2026 for fast research or deep-research-max-preview-04-2026 for maximum exhaustiveness. Agents require background=True.

Python

import time

interaction = client.interactions.create(
    agent="deep-research-preview-04-2026",
    input="Research the history of Google TPUs.",
    background=True
)
while True:
    interaction = client.interactions.get(interaction.id)
    if interaction.status == "completed":
        print(interaction.output_text)
        break
    elif interaction.status == "failed":
        print(f"Failed: {interaction.error}")
        break
    time.sleep(10)

JavaScript/TypeScript

import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

// Start background research
const initialInteraction = await client.interactions.create({
    agent: "deep-research-preview-04-2026",
    input: "Research the history of Google TPUs.",
    background: true,
});

// Poll for results
while (true) {
    const interaction = await client.interactions.get(initialInteraction.id);
    if (interaction.status === "completed") {
        console.log(interaction.output_text);
        break;
    } else if (["failed", "cancelled"].includes(interaction.status)) {
        console.log(`Failed: ${interaction.status}`);
        break;
    }
    await new Promise(resolve => setTimeout(resolve, 10000));
}

Advanced features: collaborative planning, native visualization, MCP integration, file search, multimodal inputs. See Deep Research docs.

Managed Agents

Managed agents run inside a sandboxed Linux environment hosted by Google. Fetch the Managed Agents Quickstart before writing agent code.

Antigravity Agent

The Antigravity agent (antigravity-preview-05-2026) is the general-purpose managed agent. It can execute code (Bash, Python, Node.js), manage files, browse the web, and use Google Search. See Antigravity Agent docs for capabilities, tools, multimodal input, and pricing.

Python
from google import genai

client = genai.Client()

interaction = client.interactions.create(
    agent="antigravity-preview-05-2026",
    input="Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
    environment="remote",
)

print(f"Environment ID: {interaction.environment_id}")
print(interaction.output_text)
JavaScript/TypeScript
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

const interaction = await client.interactions.create({
    agent: "antigravity-preview-05-2026",
    input: "Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
    environment: "remote",
});

console.log(`Environment ID: {interaction.environment_id}`);
console.log(interaction.output_text);

Custom Agents

See Building Custom Agents docs.

Python
agent = client.agents.create(
    id="code-reviewer",
    base_agent="antigravity-preview-05-2026",
    system_instruction="You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
    base_environment={
        "type": "remote",
        "sources": [
            {
                "type": "repository",
                "source": "https://github.com/my-org/backend",
                "target": "/workspace/repo",
            }
        ],
    },
)

# Invoke — each call forks the base environment
result = client.interactions.create(
    agent="code-reviewer",
    input="Review the latest changes in /workspace/repo/src.",
    environment="remote",
)
print(result.output_text)
JavaScript/TypeScript
const agent = await client.agents.create({
    id: "code-reviewer",
    base_agent="antigravity-preview-05-2026",
    system_instruction: "You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
    base_environment: {
        type: "remote",
        sources: [
            {
                type: "repository",
                source: "https://github.com/my-org/backend",
                target: "/workspace/repo",
            }
        ],
    },
});

const result = await client.interactions.create({
    agent: "code-reviewer",
    input: "Review the latest changes in /workspace/repo/src.",
    environment: "remote",
});
console.log(result.output_text);

Manage agents with client.agents.list(), client.agents.get(id=...), and client.agents.delete(id=...).

Streaming

Set stream=True to receive incremental server-sent events. Each stream follows: interaction.created → (step.startstep.delta(s) → step.stop)+ → interaction.completed.

Python

for event in client.interactions.create(
    model="gemini-3.5-flash",
    input="Explain quantum entanglement in simple terms.",
    stream=True,
):
    if event.event_type == "step.delta":
        if event.delta.type == "text":
            print(event.delta.text, end="", flush=True)
    elif event.event_type == "interaction.completed":
        print(f"\n\nTotal Tokens: {event.interaction.usage.total_tokens}")

JavaScript/TypeScript

const stream = await client.interactions.create({
    model: "gemini-3.5-flash",
    input: "Explain quantum entanglement in simple terms.",
    stream: true,
});
for await (const event of stream) {
    if (event.event_type === "step.delta") {
        if (event.delta.type === "text") {
            process.stdout.write(event.delta.text);
        }
    } else if (event.event_type === "interaction.completed") {
        console.log(`\n\nTotal Tokens: ${event.interaction.usage.total_tokens}`);
    }
}

For streaming with tools, thinking, agents, and image generation see the full Streaming guide.

Documentation Pages

You MUST fetch the matching page below before writing code. These hosted docs are the source of truth for parameters, types, and edge cases — do not rely solely on the examples above.

Core Documentation:

Tools & Function Calling:

Generation & Output:

Multimodal Understanding:

Files & Context:

Agents:

Advanced Features:

API Reference:

Data Model

An Interaction response contains steps, an array of typed step objects representing a structured timeline of the interaction turn.

Step Types

User steps:

  • user_input: User input (text, audio, multimodal). Contains content array.

Model/server steps:

  • model_output: Final model generation. Contains content array with text, image, audio, etc.
  • thought: Model reasoning/Chain of Thought. Has signature field (required) and optional summary.
  • function_call: Tool call request (id, name, arguments).
  • function_result: Tool result you send back (call_id, name, result).
  • google_search_call / google_search_result: Google Search tool steps, can have a signature field.
  • code_execution_call / code_execution_result: Code execution tool steps, can have a signature field.
  • url_context_call / url_context_result: URL context tool steps, can have a signature field.
  • mcp_server_tool_call / mcp_server_tool_result: Remote MCP tool steps.
  • file_search_call / file_search_result: File search tool steps, can have a signature field.

Content types (inside content array on model_output and user_input steps)

  • text: Text content (text field)
  • image / audio / document / video: Content with data, mime_type, or uri

Streaming Event Types

EventDescription
interaction.createdInteraction created; includes metadata.
interaction.status_updateInteraction-level status change.
step.startA new step begins. Contains step type and initial metadata.
step.deltaIncremental data for the current step. Contains a typed delta object.
step.stopThe step is complete. Contains index.
interaction.completedInteraction finished. Contains final usage.

Delta Types

Delta TypeParent StepDescription
textmodel_outputIncremental text token.
audiomodel_outputaudio chunk (base64).
imagemodel_outputimage chunk (base64).
thought_summarythoughtthinking summary text.
thought_signaturethoughtOpaque signature for thought verification.

Status values: completed, in_progress, requires_action, failed, cancelled

Related skills

Forks & variants (1)

Gemini Interactions Api has 1 known copy in the catalog totaling 47 installs. They canonicalize to this original listing.

FAQ

Which models should I use, and which are deprecated?

Use gemini-3.5-flash (balanced, 1M tokens), gemini-3.1-pro-preview (complex reasoning), gemini-3.1-flash-lite (cost-efficient), or gemini-3-pro-image (image generation). Models like gemini-2.5-*, gemini-2.0-*, gemini-1.5-* are deprecated. Never use legacy SDKs google-generativeai

How do I maintain conversation context across turns?

Pass previous_interaction_id from the prior response into the next create() call. Interactions are stored server-side by default (store=true): 55 days on paid tier, 1 day on free tier. Set store=false to opt out but lose state and background task support.

How do I use managed agents like Antigravity or Deep Research?

Antigravity requires environment='remote' for sandboxed code execution/web browsing. Deep Research agents require background=true and polling via interactions.get() until status='completed'. Custom agents extend base agents with system_instruction and base_environment config.

Is Gemini Interactions Api safe to install?

skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

AI & Agent Buildingllmagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.