← Back to articles
Free API Overview

Gemini API Free Tier Usage Guide

Introduction

Google AI Studio offers a free Gemini API tier — roughly 15 requests per minute and 1500 per day on the flagship Flash models — enough for personal projects and prototyping. As of 2026-08-22, Gemini officially joined the APIShare free channel pool: 10 chat models on the main ranking (gemini-3.7-flash is the current flagship, gemini-3.6-flash is Google's recommended stable pick), all with a uniform 1M-token input / 64K output envelope (the gemma-4 twins at 256K). This article demos text, multimodal, streaming, and structured output.

Architecture

flowchart LR A[AI Studio: aistudio.google.com] --> B[Create API Key] B --> C[pip install google-genai] C --> D[Call gemini-3.6-flash] D --> E[15 RPM / 1500 req/day] E --> F[1M context window]

Get an API Key

  1. Visit https://aistudio.google.com and sign in with a Google account.
  2. Click Get API Key → Create API key on the left sidebar.
  3. Copy the string starting with AIza....

Install the SDK

pip install -U google-genai

Note: use the new google-genai SDK (the old google-generativeai package is no longer evolving) — the API is cleaner.

Text Chat

import os
from google import genai

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

resp = client.models.generate_content(
    model="gemini-3.6-flash",
    contents="Explain inverted indexes in one sentence.",
)
print(resp.text)

⚠️ Thinking note: the 3.x family generates thinking chains by default and thinking tokens consume your output budget. Set max_output_tokens >= 256 or you may get empty replies that contain only thinking.

Multimodal Call

from pathlib import Path
from google import genai

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

# Image understanding
img = Path("chart.png").read_bytes()
resp = client.models.generate_content(
    model="gemini-3.6-flash",
    contents=[
        {"mime_type": "image/png", "data": img},
        "Describe the key trend of this chart in one sentence.",
    ],
)
print(resp.text)

Streaming

for chunk in client.models.generate_content_stream(
    model="gemini-3.6-flash",
    contents="Write a haiku about autumn.",
):
    print(chunk.text or "", end="", flush=True)
print()

Structured Output (JSON)

from pydantic import BaseModel
from google import genai

class Person(BaseModel):
    name: str
    title: str
    company: str

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
resp = client.models.generate_content(
    model="gemini-3.6-flash",
    contents="Extract: 'Alice is a Senior Engineer at Acme Corp.'",
    config={"response_mime_type": "response_mime_type": "application/json",
            "response_schema": Person},
)
print(resp.parsed)  # Person(name='Alice', title='Senior Engineer', company='Acme Corp.')

Safety and Quota

The free tier requires no credit card but has these limits:

  • RPM: ~15 (Flash family); lower for gemma-4 twins — check the console
  • Daily requests: ~1500 (Flash family)
  • Input tokens: 1M per minute
  • Training on your data: Free-tier data may be used by Google to improve models. Use the paid tier for production.

Model Selection

How to pick among the 10 chat models currently in the pool:

  • gemini-3.6-flash: newest flagship; first choice for complex tasks.
  • gemini-3.6-flash: Google's recommended stable pick for new projects; verified live at 21:57, first choice for high-frequency calls.
  • gemini-3.5-flash: stable generation.
  • gemini-3.1-flash-lite / -preview: lightweight pipelines and low-cost batch jobs.
  • gemini-flash-latest / gemini-flash-lite-latest: aliases that track the newest Flash/Lite if you don't want to chase version numbers.
  • gemma-4-31b-it / gemma-4-26b-a4b-it: open-source twins (256K ctx), verified working.

⚠️ Never use the gemini-2.5-* family — they return 404 for new users' API keys. Specialist tasks (image, TTS, video, music, embeddings) carry zero free-tier quota and are not in the pool — use the paid tier.

File Upload and Long Context

For large files (videos, long PDFs), first call client.files.upload() to upload, then reference the returned URI in contents — no need to re-upload on every request. The Flash family supports files up to 2GB and a 1M-token context window, ideal for whole-book summaries or long-video analysis.

Troubleshooting

  • 429 RESOURCE_EXHAUSTED: Rate limit hit. Wait 60 seconds or upgrade to the paid tier.
  • 404 model not found: You used a retired model (e.g. gemini-2.5-*). Switch to gemini-3.6-flash.
  • block_reason: SAFETY: Safety filter triggered. Adjust the prompt or lower temperature.
  • Function calling: Pass a function schema via tools; Gemini returns the call parameters.

The Gemini free tier is the cheapest way to experience multimodal AI.

Best Practices

  • 15 RPM is a hard limit: implement a local token bucket at 12 RPM to leave buffer.
  • 1M context actually caps around 800K: going beyond 800K tokens tends to hit internal limits; chunk long docs.
  • Send image bytes directly: pass Path.read_bytes() with a mime_type instead of manual base64.
  • Migrate legacy code: gemini-2.5-* returns 404 for new keys, and older gemini-1.5-* / gemini-2.0-flash have left the free-tier list — move to gemini-3.6-flash or newer.

🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →

Start on APIShare in three steps

Ready to try the options above? Three steps get you running:

  1. Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
  2. Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
  3. Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.

Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.

Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.