← Back to articles
Free API Overview

How Far Does Together AI's $5 Free Credit Go? 70B Turbo at $0.88/1M Tested: ~5M Tokens + FLUX at $0.003/Image Cost Table

Updated: 2026-08-29 · Official: https://www.together.ai · Pricing: https://www.together.ai/pricing · Console: https://api.together.ai · Verified: 2026-08-29 live pricing page + API test

One-liner: Together AI typically grants $5 free credit for new users covering text / image / embeddings in one stack, fully OpenAI-compatible at https://api.together.xyz/v1; at 70B Turbo $0.88/1M ≈ 5.68M tokens (~500 long chats) and FLUX.1-dev ~$0.003/image ≈ 1,666 images, it is the smoothest graduation path from "free trial" to "pay-as-you-go".

Why Together

  • $5 actually runs end-to-end: 5.68M tokens is enough for a prototype + a small benchmark (e.g., 1,100 long doc summaries or 7,500 short Q&As), unlike "a few hundred calls" trials that expire before you finish integration. You can validate latency, quality and cost on real workload before paying.
  • High-throughput inference cloud: Proprietary inference engine optimized for batch processing (thousands of doc summaries / embeddings). Cost per 1K docs is lower than low-latency providers when you can tolerate 1-2s per request.
  • One billing for all modalities: Chat / Image / Embeddings / Rerank / Moderation share the same $5 pool. A full RAG stack (embed + retrieve + generate + image) needs no separate bills or keys.
  • Zero migration: OpenAI-compatible endpoint — just change base_url to https://api.together.xyz/v1. Dashboard has a "Pause when credits run out" switch to prevent accidental charges after $5 is exhausted.

Free API Deep Dive (Verified 2026-08-29)

Model API ID Capability Context Pricing ($ / 1M tokens) What $5 Gets You
Llama-3.3-70B Turbo meta-llama/Llama-3.3-70B-Instruct-Turbo Balanced / General 128K $0.88 in / $0.88 out ~5.68M tokens ≈ 500 chats (1K in + 1K out each)
Qwen2.5-7B Turbo Qwen/Qwen2.5-7B-Instruct-Turbo Bilingual EN/ZH 32K $0.18 / 1M ~27.7M tokens ≈ 2,500 chats
DeepSeek-R1 deepseek-ai/DeepSeek-R1 Reasoning 64K $3.00 in / $7.00 out ~0.71M tokens ≈ 70 long reasoning chats
FLUX.1-dev black-forest-labs/FLUX.1-dev Text-to-Image — ~$0.003 / image (1024×1024, 28 steps) ~1,666 images
FLUX.1-schnell black-forest-labs/FLUX.1-schnell Text-to-Image (fast) — ~$0.0015 / image ~3,333 images
BGE-M3 Embedding BAAI/bge-m3 Embeddings 8K $0.02 / 1M ~250M tokens

Source: https://www.together.ai/pricing snapshot 2026-08-29; $5 is the typical new-user grant, subject to console at signup. R1 is expensive on $5 — reserve it for hard reasoning only.

Cost Breakdown (at 70B Turbo $0.88/1M)

Scenario Tokens per Call Cost per Call How Many on $5 Notes
Short Q&A (500 in + 250 out) 750 ~$0.00066 ~7,500 Customer support / FAQ
Long chat (1K in + 500 out) 1,500 ~$0.00132 ~3,700 RAG summarization
Batch summary (4K in + 1K out) 5,000 ~$0.0044 ~1,100 docs Doc batch processing
FLUX 1024×1024 — $0.003 / image ~1,666 images Same $5 pool as text
Mixed: 300 chats (0.6M, $0.53) + 200 images ($0.60) — — still leaves $3.87 Text+image in one pool

Rate Limits & Quota (Verified via Dashboard)

Dimension Value Details
Free credit $5 One-time for new users, expiry per console (typically 30-90 days)
Rate Dynamic per model, 70B ~30-60 RPM Visible in Dashboard → Metrics, throttles with 429 on exceed
After exhaustion Auto pay-as-you-go if card bound Enable "Pause when credits run out" to stop instead
Observability Dashboard → Usage / Metrics Per-model latency, throughput, spend breakdown

5-Min Quick Start

1) Signup & API Key

  1. Sign up at https://api.together.ai/signup
  2. Create key at https://api.together.ai/settings/api-keys → tgp_...
  3. Confirm $5 balance and toggle Pause when credits run out in Dashboard → Billing

2) One-Click Calls (curl / Python / Node.js — all verified HTTP 200)

curl — Chat (70B Turbo)

curl -X POST https://api.together.xyz/v1/chat/completions \
  -H "Authorization: Bearer $TOGETHER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo",
    "messages": [{"role":"user","content":"Write a one-line product slogan"}],
    "max_tokens": 100
  }'
# Expected: HTTP 200, body contains choices[0].message.content and usage.prompt_tokens / completion_tokens

curl — Image (FLUX.1-dev)

curl -X POST https://api.together.xyz/v1/images/generations \
  -H "Authorization: Bearer $TOGETHER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "black-forest-labs/FLUX.1-dev",
    "prompt": "a cyberpunk street at night, neon lights, ultra detailed",
    "width": 1024,
    "height": 1024,
    "steps": 28
  }'
# Expected: HTTP 200, data[0].url is the image URL

Python (OpenAI SDK + requests for image)

from openai import OpenAI
import requests

client = OpenAI(base_url="https://api.together.xyz/v1", api_key="tgp_...")

# Chat
resp = client.chat.completions.create(
    model="meta-llama/Llama-3.3-70B-Instruct-Turbo",
    messages=[{"role": "user", "content": "Introduce Together AI in one sentence"}],
    max_tokens=100
)
print(resp.choices[0].message.content)
print(resp.usage)  # prompt_tokens / completion_tokens / total_tokens

# Image
r = requests.post(
    "https://api.together.xyz/v1/images/generations",
    headers={"Authorization": f"Bearer {client.api_key}", "Content-Type": "application/json"},
    json={"model": "black-forest-labs/FLUX.1-dev", "prompt": "a cute cat astronaut", "width": 1024, "height": 1024, "steps": 28}
)
print(r.json()["data"][0]["url"])

Node.js (OpenAI SDK)

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.together.xyz/v1",
  apiKey: process.env.TOGETHER_API_KEY
});

// Chat
const r = await client.chat.completions.create({
  model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
  messages: [{ role: "user", content: "Hello from Together" }],
  max_tokens: 100
});
console.log(r.choices[0].message.content, r.usage);

// For image, use fetch to /v1/images/generations with same baseURL

3) Practical Tips

  • Start with Qwen2.5-7B, then upgrade to 70B: 7B is 5× cheaper (27.7M tokens on $5) and handles high-frequency calls, bilingual and simple summarization well. Reserve 70B/R1 for hard tasks that need quality/reasoning.
  • Batch workloads belong on Together: For 1K+ doc summaries or embeddings, Together's throughput pricing beats low-latency providers. Monitor per-model p50 latency in Dashboard before locking in 70B vs 7B.
  • Always enable "Pause when credits run out": In Dashboard → Billing, enable Pause before binding a card. Otherwise $5 exhaustion silently becomes pay-as-you-go.

Pricing & Pitfalls

Pitfall Symptom Fix
$5 expiry Balance drops to 0 after 30-90 days Use within 30 days of signup, check expiry in Billing
Shared pool Image generation burns chat budget Pre-allocate: e.g., $3 for chat + $2 for images, watch Usage daily
70B latency/cost Small tasks also routed to 70B, slow & expensive Default to Qwen2.5-7B, switch to 70B/R1 only for hard prompts
Pause not enabled Auto-charged after $5 Enable Pause in Billing; set budget alert at $4

Graduation path: After $5 is spent, Together's pay-as-you-go remains cheaper than closed-source (70B $0.88/1M vs GPT-4o $5/1M input). It is the cheapest first paid step. If you need pure free, fall back to OpenRouter :free models (e.g., qwen/qwen2.5-7b-instruct:free) as a safety net — same OpenAI-compatible code, just swap base_url and model ID.

Official Resources (Traceability)

Verified 2026-08-29 pricing and model list; article will be updated within 24h on price change.


🚀 Get Started: One-Click Free API Access

Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.

👉 Register on Apishare.cc → Get your unified API Key

📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →

Start on APIShare in three steps

Ready to try the options above? Three steps get you running:

  1. Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
  2. Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
  3. Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.

Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.

Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.


About the Free API Aggregator

The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.

More in this category

Free API Cost and Quota Control in Practice: 429 Backoff, RPM Budgets, and Multi-Model Fallback2026 Free OneAPI Unified Gateway: Connect 100+ LLM APIs at Zero Cost in One GuideRun a 550B-Parameter Model for Free: 2026 Nemotron 3 Ultra Complete Guide (OpenRouter Free Tier Tested)2026 Free Embedding Vector Model API Panorama: BGE-M3 / Voyage / Nomic / Google / Azure and 6 Options Tested (September Update)Free Function Calling / Tool Use API Tutorial: DeepSeek / Gemini / Qwen — Zero-Cost Agent Tooling (2026-09-16 Verified)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.