Updated: 2026-08-29 · Official: https://www.together.ai · Pricing: https://www.together.ai/pricing · Console: https://api.together.ai · Verified: 2026-08-29 live pricing page + API test
One-liner: Together AI typically grants $5 free credit for new users covering text / image / embeddings in one stack, fully OpenAI-compatible at
https://api.together.xyz/v1; at 70B Turbo $0.88/1M ≈ 5.68M tokens (~500 long chats) and FLUX.1-dev ~$0.003/image ≈ 1,666 images, it is the smoothest graduation path from "free trial" to "pay-as-you-go".
Why Together
- $5 actually runs end-to-end: 5.68M tokens is enough for a prototype + a small benchmark (e.g., 1,100 long doc summaries or 7,500 short Q&As), unlike "a few hundred calls" trials that expire before you finish integration. You can validate latency, quality and cost on real workload before paying.
- High-throughput inference cloud: Proprietary inference engine optimized for batch processing (thousands of doc summaries / embeddings). Cost per 1K docs is lower than low-latency providers when you can tolerate 1-2s per request.
- One billing for all modalities: Chat / Image / Embeddings / Rerank / Moderation share the same $5 pool. A full RAG stack (embed + retrieve + generate + image) needs no separate bills or keys.
- Zero migration: OpenAI-compatible endpoint — just change
base_urltohttps://api.together.xyz/v1. Dashboard has a "Pause when credits run out" switch to prevent accidental charges after $5 is exhausted.
Free API Deep Dive (Verified 2026-08-29)
| Model | API ID | Capability | Context | Pricing ($ / 1M tokens) | What $5 Gets You |
|---|---|---|---|---|---|
| Llama-3.3-70B Turbo | meta-llama/Llama-3.3-70B-Instruct-Turbo |
Balanced / General | 128K | $0.88 in / $0.88 out | ~5.68M tokens ≈ 500 chats (1K in + 1K out each) |
| Qwen2.5-7B Turbo | Qwen/Qwen2.5-7B-Instruct-Turbo |
Bilingual EN/ZH | 32K | $0.18 / 1M | ~27.7M tokens ≈ 2,500 chats |
| DeepSeek-R1 | deepseek-ai/DeepSeek-R1 |
Reasoning | 64K | $3.00 in / $7.00 out | ~0.71M tokens ≈ 70 long reasoning chats |
| FLUX.1-dev | black-forest-labs/FLUX.1-dev |
Text-to-Image | — | ~$0.003 / image (1024×1024, 28 steps) | ~1,666 images |
| FLUX.1-schnell | black-forest-labs/FLUX.1-schnell |
Text-to-Image (fast) | — | ~$0.0015 / image | ~3,333 images |
| BGE-M3 Embedding | BAAI/bge-m3 |
Embeddings | 8K | $0.02 / 1M | ~250M tokens |
Source: https://www.together.ai/pricing snapshot 2026-08-29; $5 is the typical new-user grant, subject to console at signup. R1 is expensive on $5 — reserve it for hard reasoning only.
Cost Breakdown (at 70B Turbo $0.88/1M)
| Scenario | Tokens per Call | Cost per Call | How Many on $5 | Notes |
|---|---|---|---|---|
| Short Q&A (500 in + 250 out) | 750 | ~$0.00066 | ~7,500 | Customer support / FAQ |
| Long chat (1K in + 500 out) | 1,500 | ~$0.00132 | ~3,700 | RAG summarization |
| Batch summary (4K in + 1K out) | 5,000 | ~$0.0044 | ~1,100 docs | Doc batch processing |
| FLUX 1024×1024 | — | $0.003 / image | ~1,666 images | Same $5 pool as text |
| Mixed: 300 chats (0.6M, $0.53) + 200 images ($0.60) | — | — | still leaves $3.87 | Text+image in one pool |
Rate Limits & Quota (Verified via Dashboard)
| Dimension | Value | Details |
|---|---|---|
| Free credit | $5 | One-time for new users, expiry per console (typically 30-90 days) |
| Rate | Dynamic per model, 70B ~30-60 RPM | Visible in Dashboard → Metrics, throttles with 429 on exceed |
| After exhaustion | Auto pay-as-you-go if card bound | Enable "Pause when credits run out" to stop instead |
| Observability | Dashboard → Usage / Metrics | Per-model latency, throughput, spend breakdown |
5-Min Quick Start
1) Signup & API Key
- Sign up at https://api.together.ai/signup
- Create key at https://api.together.ai/settings/api-keys →
tgp_... - Confirm $5 balance and toggle Pause when credits run out in Dashboard → Billing
2) One-Click Calls (curl / Python / Node.js — all verified HTTP 200)
curl — Chat (70B Turbo)
curl -X POST https://api.together.xyz/v1/chat/completions \
-H "Authorization: Bearer $TOGETHER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.3-70B-Instruct-Turbo",
"messages": [{"role":"user","content":"Write a one-line product slogan"}],
"max_tokens": 100
}'
# Expected: HTTP 200, body contains choices[0].message.content and usage.prompt_tokens / completion_tokens
curl — Image (FLUX.1-dev)
curl -X POST https://api.together.xyz/v1/images/generations \
-H "Authorization: Bearer $TOGETHER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "black-forest-labs/FLUX.1-dev",
"prompt": "a cyberpunk street at night, neon lights, ultra detailed",
"width": 1024,
"height": 1024,
"steps": 28
}'
# Expected: HTTP 200, data[0].url is the image URL
Python (OpenAI SDK + requests for image)
from openai import OpenAI
import requests
client = OpenAI(base_url="https://api.together.xyz/v1", api_key="tgp_...")
# Chat
resp = client.chat.completions.create(
model="meta-llama/Llama-3.3-70B-Instruct-Turbo",
messages=[{"role": "user", "content": "Introduce Together AI in one sentence"}],
max_tokens=100
)
print(resp.choices[0].message.content)
print(resp.usage) # prompt_tokens / completion_tokens / total_tokens
# Image
r = requests.post(
"https://api.together.xyz/v1/images/generations",
headers={"Authorization": f"Bearer {client.api_key}", "Content-Type": "application/json"},
json={"model": "black-forest-labs/FLUX.1-dev", "prompt": "a cute cat astronaut", "width": 1024, "height": 1024, "steps": 28}
)
print(r.json()["data"][0]["url"])
Node.js (OpenAI SDK)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.together.xyz/v1",
apiKey: process.env.TOGETHER_API_KEY
});
// Chat
const r = await client.chat.completions.create({
model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
messages: [{ role: "user", content: "Hello from Together" }],
max_tokens: 100
});
console.log(r.choices[0].message.content, r.usage);
// For image, use fetch to /v1/images/generations with same baseURL
3) Practical Tips
- Start with Qwen2.5-7B, then upgrade to 70B: 7B is 5× cheaper (27.7M tokens on $5) and handles high-frequency calls, bilingual and simple summarization well. Reserve 70B/R1 for hard tasks that need quality/reasoning.
- Batch workloads belong on Together: For 1K+ doc summaries or embeddings, Together's throughput pricing beats low-latency providers. Monitor per-model p50 latency in Dashboard before locking in 70B vs 7B.
- Always enable "Pause when credits run out": In Dashboard → Billing, enable Pause before binding a card. Otherwise $5 exhaustion silently becomes pay-as-you-go.
Pricing & Pitfalls
| Pitfall | Symptom | Fix |
|---|---|---|
| $5 expiry | Balance drops to 0 after 30-90 days | Use within 30 days of signup, check expiry in Billing |
| Shared pool | Image generation burns chat budget | Pre-allocate: e.g., $3 for chat + $2 for images, watch Usage daily |
| 70B latency/cost | Small tasks also routed to 70B, slow & expensive | Default to Qwen2.5-7B, switch to 70B/R1 only for hard prompts |
| Pause not enabled | Auto-charged after $5 | Enable Pause in Billing; set budget alert at $4 |
Graduation path: After $5 is spent, Together's pay-as-you-go remains cheaper than closed-source (70B $0.88/1M vs GPT-4o $5/1M input). It is the cheapest first paid step. If you need pure free, fall back to OpenRouter :free models (e.g., qwen/qwen2.5-7b-instruct:free) as a safety net — same OpenAI-compatible code, just swap base_url and model ID.
Official Resources (Traceability)
- Official: https://www.together.ai
- Pricing: https://www.together.ai/pricing
- Docs: https://docs.together.ai
- Console / Keys: https://api.together.ai
- Status: https://status.together.ai
Verified 2026-08-29 pricing and model list; article will be updated within 24h on price change.
🚀 Get Started: One-Click Free API Access
Want to call all the free models above with a single API key, no need to sign up for each provider? Apishare.cc provides a unified API Key — one key, 100+ models, free models at zero cost.
👉 Register on Apishare.cc → Get your unified API Key
📊 Want to see more free model rankings? Check out the Sep 2026 Free LLM API Rankings →
Start on APIShare in three steps
Ready to try the options above? Three steps get you running:
- Create an account - open the APIShare free API registration page. An email address is all you need; no credit card required.
- Browse the free API catalog - head to the complete free API list and filter by text, image, audio, embedding, or multimodal. Each entry shows its free quota, rate limit, and availability status.
- Grab a key and integrate - generate an API key in your dashboard and paste it into your application. Every plan includes actively-updated APIs gateways covering every provider mentioned in this guide.
Already have an account? Use the APIShare login page, or visit the APIShare homepage for a full platform overview. Registration is free, and you can stop at any time.
Every outbound link in this guide carries a UTM parameter (utm_source=apishare_devto&utm_medium=referral&utm_campaign=free_api_article) for clean campaign attribution.
About the Free API Aggregator
The models covered in this guide are all served through the APIShare free API aggregator, which gives you one key for the whole catalog.
- Full model catalog: APIShare free API directory
- Sign up for a free trial key: Register and claim your API key