Gemini API
发布时间:2026-09-11 | 浏览:1
Español – América Latina
Português – Brasil
The Gemini API is the fastest path from prompt to production with Gemini, Veo, Nano Banana, and more. It lets you integrate these generative models into your applications to generate text and images, analyze multimodal inputs, and build conversational agents.
Meet the models
spark Gemini 3.8 Flash New
Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
spark Gemini 3.5 Flash-Lite
High-volume, cost-sensitive model optimized for low-latency high throughput subagent tasks.
auto_awesome Gemini 3.1 Pro
Our most intelligent model, the best in the world for multimodal understanding, all built on state-of-the-art reasoning.
🍌 Nano Banana 2 and Nano Banana Pro
State-of-the-art image generation and editing models.
video_library Gemini Omni Flash
Our state-of-the-art video generation and editing model.
speech_to_text Gemini 3.5 Transcribe New
Low-latency speech-to-text model with utterance-based language detection, speaker diarization, and word timestamps.
spark Gemini Robotics
A vision-language model (VLM) that brings Gemini's agentic capabilities to robotics and enables advanced reasoning in the physical world.
Explore Capabilities
Image Generation
Generate and edit highly contextual images natively with Nano Banana.
Input millions of tokens to Gemini models and derive understanding from unstructured images, videos, and documents.
Structured Outputs
Constrain Gemini to respond with JSON, a structured data format suitable for automated processing.
Function Calling
Build agentic workflows by connecting Gemini to external APIs and tools.
Video Generation with Veo 3.1
Create high-quality video content from text or image prompts with our state-of-the-art model.
Voice Agents with Live API
Build real-time voice applications and agents with the Live API.
Connect Gemini to the world through built-in tools like Google Search, URL Context, Google Maps, Code Execution and Computer Use.
Document Understanding
Process up to 1000 pages of PDF files with full multimodal understanding or other text-based file types.
Explore how thinking capabilities improve reasoning for complex tasks and agents.
Interactions API
The Interactions API has become our default interface as of June 2026 and is the best way to build with Gemini models and agents going forward. If you're starting a new project, you should use the Interactions API. While it remains supported, the generateContent API is now considered legacy.
Interactions Overview
Learn how the Interactions API manages conversation state, messages, and output formats.
Migration Guide
Step-by-step guide to transition your code from generateContent to the Interactions API.
Stream real-time tokens, incremental thoughts, and tool call events.
Test prompts, manage your API keys, monitor usage, and build prototypes.
Ask questions and find solutions from other developers and Google engineers.
Find detailed information about the Gemini API in the official reference documentation.
Check the status of Gemini API, Google AI Studio, and our model services.
Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License , and code samples are licensed under the Apache 2.0 License . For details, see the Google Developers Site Policies . Java is a registered trademark of Oracle and/or its affiliates.
Last updated 2026-09-10 UTC.