Free Code Completion APIs in Practice: FIM Formatting, Context Trimming, and Latency Tuning
Most people who try code completion on a free tier hit the same wall on the very first attempt. They paste the whole conversation into the model and ask it to "continue." The result is a completion that lands in the wrong place, latency that climbs past two seconds, and the code sitting after the cursor getting swallowed as if it were prior context. This article explains three things properly: how to assemble the FIM prompt, how to trim the context window, and how to push latency down to something usable. Every model referenced here is one that is already verified as available on the free tier.
1. Two Different Things Both Called "Code Completion"
Inline completion and conversational coding are two distinct tasks. Using the wrong endpoint for the job is the number one cause of both latency blowups and misplaced completions.
| Dimension | Inline Completion | Conversational Coding |
|---|---|---|
| Trigger | Fires automatically as you type, possibly several times per second | Fires when the user explicitly asks |
| Input shape | Code before the cursor + code after the cursor | A natural language question + a code snippet |
| Expected output | A middle fill of a few to a few dozen tokens | A full explanation, a whole function, multi-turn discussion |
| Latency budget | Under 500ms, otherwise typing feels stuck | 2-5 seconds is acceptable |
| Correct endpoint | FIM / Completions | Chat Completions |
| Typical free models | mistral-code-fim-latest |
qwen3-coder-30b, deepseek-coder |
The key conclusion: inline completion must go through the FIM format. If you drive inline completion with a Chat endpoint, the model treats your "code after the cursor" as part of the conversation history and then produces a brand new block of code that does not connect to anything. That is exactly where "the completion lands in the wrong place" comes from.
2. FIM Format: Assembling the Three Special Tokens Correctly
FIM stands for Fill-In-the-Middle. The whole idea is to split the source into three segments and use special markers to tell the model "fill the gap in the center."
2.1 Two Orderings: PSM and SPM
| Ordering | Structure | Where it applies |
|---|---|---|
| PSM (Prefix-Suffix-Middle) | prefix → suffix → middle | The mainstream choice, default for the Mistral family |
| SPM (Suffix-Prefix-Middle) | suffix → prefix → middle | Certain StarCoder variants |
The PSM assembly, using Mistral style markers as the example:
| Segment | Marker | Content |
|---|---|---|
| Prefix | <|fim_prefix|> |
Everything before the cursor |
| Suffix | <|fim_suffix|> |
Everything after the cursor |
| Middle | <|fim_middle|> |
Left empty; generation starts here |
2.2 The Three Most Common Assembly Mistakes
- Do not reorder the segments. Writing
prefix → middle → suffixis the single most frequent error. During training the model only ever sawprefix → suffix → middle. Get the order wrong and completion quality falls off a cliff. - Do not omit the suffix. A lot of people send only the prefix, which silently degrades FIM into plain continuation. The suffix is the entire value of FIM over continuation — it tells the model what already exists downstream, so the generated variable names and indentation actually line up.
- The markers must be byte exact. The pipe inside
<|fim_prefix|>is the ASCII pipe|(U+007C), not the full width|. Copy pasting from a Chinese language document is a reliable way to get it swapped by an input method. Once swapped, the model does not recognize the marker at all and treats it as ordinary text.
3. Context Trimming: The Life or Death Line on a Free Tier
Free tiers usually ship a context window far smaller than the paid one. Your trimming strategy decides whether the feature runs at all.
3.1 Retention Priority, Highest First
| Priority | Content | Retention rule |
|---|---|---|
| 1 | The 50 lines before the cursor | Must be kept in full |
| 2 | The 20 lines after the cursor | Must be kept; this is the FIM suffix |
| 3 | The import block of the current file | Pull it back in even if it sits outside those 50 lines |
| 4 | Function signatures from other files in the project | Add only if budget remains |
| 5 | Historical code 50-200 lines above the cursor | First thing to drop when budget is tight |
3.2 Trimming by Token Budget, Not by Line Count
Never trim by line count. Trim by token count. Chinese comments and English code differ in token density by more than three times, so a line based rule will overshoot or undershoot unpredictably.
- Set a total budget
B. For an 8K free window, reserve 1K for output, soB = 7000. - Place the suffix first; call its cost
stokens. - Then add the last N lines of the prefix, walking backwards from the cursor, until the running total approaches
B - s. - If budget remains, insert the import block at the very front of the prefix.
- If you are over budget, drop from the head of the prefix, never from the tail. Code closer to the cursor matters more.
3.3 Indentation Alignment: The Invisible Killer
If trimming leaves the last line of the prefix as a half statement — say you cut right through if (a && — the model will generate syntactically broken output. The rule is simple: trim boundaries must land on complete lines. Dropping one extra line is always better than keeping half of one.
4. Latency Tuning: From Two Seconds Down to 400ms
The latency budget for inline completion is brutal. People type at roughly 300ms per character, so anything above 500ms is perceived as lag.
4.1 Five Optimizations You Can Ship Today
| Optimization | What to do | Expected gain |
|---|---|---|
| Cap max_tokens | Set 32-64 for inline completion; never leave the default | Largest single gain, often saves 60%+ |
| Stop sequences | Set stop=["\n\n", "```"] to halt at a blank line |
Prevents generating an entire function |
| Debounce | Wait 250ms after typing stops before sending | Cuts roughly 70% of useless requests |
| Request cancellation | Abort the in-flight request when a new keystroke arrives | Stops stale results overwriting the new position |
| Streaming | Use stream=True and render on the first token |
Perceived latency drops under 200ms |
4.2 Why max_tokens Is Priority Number One
Generation time scales close to linearly with output token count. Defaults are frequently 1024 or 2048, while inline completion actually needs 10 to 30 tokens. Dropping max_tokens from 1024 to 48 is a theoretical 20x difference. In practice, because time to first token is a fixed cost, end to end you typically save 60-80%.
4.3 Realistic Expectations on a Free Tier
| Scenario | Realistic latency band | Usable for inline completion? |
|---|---|---|
| Free tier + small model (7B class) | 300-800ms | Yes, with debouncing |
| Free tier + large model (70B class) | 1.5-4s | No; conversational use only |
| Free tier at peak hours | Fluctuates 2-5x | Timeout plus silent fallback is mandatory |
Conclusion: small models for inline completion, large models for conversational coding. Driving inline completion with something like qwen3-coder-30b will never meet the latency bar. That is not a configuration problem, it is physics.
5. A Runnable Example: One Complete FIM Call
This is the only place in the article that needs a code block, because the assembly order of the FIM prompt is the core teaching content.
import os
from openai import OpenAI
# Unified free tier entry point; grab the key from the console in one click
client = OpenAI(
api_key=os.environ["APISHARE_KEY"],
base_url="https://apishare.cc/v1",
)
def build_fim_prompt(code: str, cursor: int) -> str:
"""Split source into PSM segments at the cursor position."""
prefix, suffix = code[:cursor], code[cursor:]
# Trim boundaries must land on complete lines, never mid statement
lines = prefix.split("\n")
kept = lines[-50:] if len(lines) > 50 else lines
prefix = "\n".join(kept)
suffix = "\n".join(suffix.split("\n")[:20])
return f"<|fim_prefix|>{prefix}<|fim_suffix|>{suffix}<|fim_middle|>"
def complete(code: str, cursor: int) -> str:
resp = client.completions.create(
model="mistral-code-fim-latest",
prompt=build_fim_prompt(code, cursor),
max_tokens=48, # the make or break parameter
temperature=0.1, # completion wants determinism, not creativity
stop=["\n\n", "```"],
stream=False,
)
return resp.choices[0].text
if __name__ == "__main__":
src = "def total(items):\n s = 0\n for it in items:\n \n return s\n"
pos = src.index(" \n") + 8
print(repr(complete(src, pos)))
Why each of those four parameters is set that way:
| Parameter | Value | Reason |
|---|---|---|
max_tokens |
48 | Inline completion almost never exceeds 30 tokens; 48 leaves headroom without waste |
temperature |
0.1 | Completion needs determinism; high temperature makes variable names drift randomly |
stop |
["\n\n", "```"] |
A blank line means the logical block ended; continuing is out of bounds |
model |
A FIM specific model | Ordinary chat models were never trained on FIM markers, so the markers do nothing |
6. Capability Radar: The Tradeoffs Between Three Model Classes
How to read this: there is no all round winner. The FIM specific small model dominates on latency and format support but is weak at repository level understanding. The Coder large model is the mirror image. The correct engineering answer is to wire up both — inline completion goes to the small model, while an explicit Ctrl+K question goes to the large one.
7. Troubleshooting Table
| Symptom | Most likely cause | Diagnostic action |
|---|---|---|
| Completion duplicates code after the cursor | Suffix was not sent; degraded into continuation | Print the actual prompt and confirm there is content after <|fim_suffix|> |
| Completion lands in the wrong place | Chat endpoint used instead of Completions | Switch to the FIM endpoint and verify the model is FIM capable |
| Generates a huge unrelated block | max_tokens left at default |
Drop it to 48 and set stop |
| Indentation is a mess | Trim cut through a half statement | Align trim boundaries to complete lines |
| Markers treated as plain text | Pipe swapped to full width | by an input method |
Use ASCII |, copy from code rather than from docs |
| 401 / 403 | Model name pasted as a hash, or namespace prefix missing | See the troubleshooting article on this site; model names must carry the namespace |
| Mass timeouts at peak | Free tier rate limiting | Add an 800ms timeout plus silent fallback; never block the editor |
8. Pre Launch Acceptance Checklist
- Format: the three markers appear in prefix → suffix → middle order, and the pipes are ASCII
- Trimming: boundaries land on complete lines; prefix at most 50 lines, suffix at most 20; the import block has been pulled back in
- Parameters:
max_tokensat most 64,temperatureat most 0.2,stopconfigured - Latency: end to end P95 under 500ms on a small model; timeout and silent fallback in place
- Interaction: 250ms debounce; a new keystroke aborts the previous request; first token renders via streaming
- Quota: daily request count tracked, with automatic switch to a backup model as the free ceiling approaches
9. Going Deeper on Repository Level Context
Everything above assumes a single file. Real projects span dozens of files, and that is where the second half of the engineering work lives.
9.1 Signature Indexing Beats Full File Injection
A common mistake is stuffing entire related files into the prompt. That burns the budget instantly and mostly adds noise. A far better approach on a constrained free tier is to index signatures only:
| What to index | Why it pays off |
|---|---|
| Function and method signatures | The model needs names and parameter shapes, not bodies |
| Class declarations with base classes | Reveals the inheritance contract cheaply |
| Exported constants and enum members | Prevents the model inventing plausible but wrong names |
| Type aliases and interface definitions | Keeps generated annotations consistent |
A signature index for a medium project typically costs 500 to 1500 tokens, versus 10,000+ for the same files in full. That difference is the entire reason repository aware completion is feasible on a free tier at all.
9.2 Retrieval Order Matters More Than Retrieval Volume
When you do pull in cross file context, order it by relevance and place the most relevant chunk closest to the prefix boundary. Models weight nearby tokens more heavily. A common failure is alphabetical ordering, which puts the least relevant file right where it does the most damage.
9.3 Cache the Index, Not the Completion
The signature index changes only when files change. Cache it aggressively and invalidate on save. Completions, by contrast, should never be cached — the cursor position differs every time, and a stale completion is worse than no completion.
10. Handling Free Tier Rate Limits Gracefully
Free tiers enforce rate limits, and inline completion is the most request hungry feature in any editor. Getting this wrong means the feature works in a demo and fails in real use.
10.1 A Three Tier Degradation Ladder
| Tier | Condition | Behavior |
|---|---|---|
| Normal | Under quota, latency healthy | Full FIM completion, streaming enabled |
| Throttled | Approaching quota or latency above 1.2s | Raise debounce to 500ms, cut max_tokens to 24 |
| Silent off | Quota exhausted or repeated 429s | Stop sending requests entirely; show no error to the user |
The critical design decision is the third tier. Never surface a rate limit error inside an editor. The user is typing; an error popup mid keystroke is far more disruptive than the absence of a suggestion. Degrade silently and resume automatically when the window resets.
10.2 Counting Requests Correctly
Debounce and cancellation together typically reduce request volume by 70% or more. If you are still hitting limits, the next lever is trigger policy: fire only after a word boundary or an opening bracket, rather than on every single character. That single change often halves the request count with no perceptible loss in usefulness.
11. Evaluating Completion Quality Without Guessing
"Feels good" is not a metric. Three cheap measurements will tell you whether a configuration change actually helped.
| Metric | Definition | Target on a free tier |
|---|---|---|
| Acceptance rate | Accepted completions divided by shown completions | Above 25% is healthy |
| Prefix match rate | Generated text that exactly continues the prefix indentation | Above 90% |
| P95 latency | 95th percentile end to end time | Under 500ms for inline |
Acceptance rate is the one that matters. If it sits below 15%, the problem is almost never the model — it is your trimming or your max_tokens. Fix those first, then consider switching models.
13. Language Specific Behavior You Cannot Ignore
A single set of trimming and parameter defaults will not serve every language well. Token density, comment style, and indentation semantics all differ, and those differences show up directly in completion quality.
| Language family | Token density | Practical consequence |
|---|---|---|
| Python | Low, whitespace significant | Indentation errors are fatal; a single misaligned space breaks the block |
| JavaScript / TypeScript | Medium | Semicolon insertion and async patterns confuse models trained mostly on synchronous code |
| Go | Medium | Strict formatting means the model must emit gofmt compatible output or the diff looks noisy |
| Java / C# | High, verbose | Boilerplate dominates; the 50 line prefix window often contains only imports and class declarations |
| Rust | High | Borrow checker constraints mean a syntactically valid completion may still not compile |
13.1 Adjusting the Prefix Window per Language
For verbose languages such as Java, a fixed 50 line prefix frequently contains no executable logic at all — only package declarations, imports, and annotations. The fix is to make the window semantic rather than purely positional: always include the enclosing method signature even when it sits above the 50 line cutoff. For Python, the opposite adjustment applies — the enclosing function definition plus its decorators is usually enough, and pulling in more tends to add unrelated sibling functions.
13.2 Comment Language Affects Token Cost
Comments written in Chinese consume roughly three times the tokens of equivalent English comments for the same information. On a tight free tier budget this is not a trivial detail. If your codebase has heavy Chinese comments, consider stripping comment lines from the prefix beyond the nearest 10 lines. The model rarely needs a distant comment to predict the next statement, and the token savings can be the difference between fitting the import block and not.
14. Building a Regression Test Set for Completion
Completion behavior changes when you switch models, adjust trimming, or when the provider silently updates a checkpoint. Without a test set you will not notice a regression until users complain.
14.1 What a Minimal Test Set Looks Like
Twenty cases are enough to catch most regressions. Build them from real code in your own repository rather than from synthetic examples.
| Case category | Count | What it verifies |
|---|---|---|
| Function body completion | 5 | Basic continuation and variable naming |
| Argument list completion | 4 | Signature awareness and type correctness |
| Conditional branch completion | 3 | Control flow and indentation |
| Import statement completion | 3 | Whether the model invents nonexistent modules |
| Cross file reference completion | 3 | Whether your signature index is actually being used |
| Empty context edge case | 2 | Behavior when the prefix is trivially short |
14.2 Assertions That Do Not Require an LLM Judge
Do not reach for model based scoring when deterministic checks will do. For each case assert:
- The output is non empty and contains no FIM marker leakage
- The output contains no triple backticks, which would mean it escaped the stop sequence
- Indentation of the first generated line matches the expected column exactly
- The concatenation of prefix plus generated text parses without a syntax error, checked with the language native parser
- End to end latency is under the configured ceiling
Assertion four is the most valuable one in the entire set. A completion that produces a syntax error when concatenated is useless regardless of how plausible it reads, and this check catches it every time at zero cost.
14.3 Running It in CI Without Burning Quota
Free tier quota is limited, so running twenty cases on every commit is not viable. Run the full set nightly, and on pull requests run only the five function body cases. Record the acceptance metrics over time; a sudden drop in the nightly run is your early warning that the provider changed something.
15. Migration Path: From Chat Based to FIM Based Completion
If you already shipped completion using a Chat endpoint, migrating to FIM is not a rewrite. The change is localized and can be done incrementally.
| Step | Action | Risk |
|---|---|---|
| 1 | Add a feature flag that routes inline completion to the FIM path | None; default stays on the old path |
| 2 | Implement build_fim_prompt and unit test the marker order in isolation |
Low |
| 3 | Verify the target model actually supports FIM markers with a single manual call | Low |
| 4 | Enable the flag for internal users only, measure acceptance rate for three days | Low |
| 5 | Compare acceptance rate and P95 latency against the Chat baseline | None |
| 6 | Roll out to all users, keep the Chat path as the automatic fallback | Low |
Step six deserves emphasis. Keep the Chat based path alive as a fallback rather than deleting it. FIM specific models are a smaller pool than chat models, so when the FIM model is rate limited or temporarily unavailable, falling back to the chat path with a degraded but functional completion is better than showing nothing at all.
15.1 What Typically Improves, and by How Much
Based on the mechanics described above rather than on any vendor claim, the expected direction of change is consistent:
| Metric | Direction after migration | Why |
|---|---|---|
| Acceptance rate | Up | The suffix now constrains generation to something that actually connects |
| Misplaced completions | Down sharply | The model no longer treats downstream code as history |
| Output length variance | Down | max_tokens plus stop sequences bound the response |
| P95 latency | Down | Smaller models and capped output both reduce generation time |
| Token cost per request | Down | Capped output dominates the cost calculation |
The one metric that can move the wrong way is coverage. A FIM specific small model may know fewer languages than a large general chat model. That is precisely why the fallback in step six matters, and why the language breakdown in section thirteen should inform your model choice rather than being discovered afterwards.
16. A Complete Worked Session
To make the whole picture concrete, walk through one realistic editing session end to end.
The user is editing a Python utility module. They type def normalize(rows): and press enter. The editor inserts a fresh indented line and the cursor sits at column four.
- Debounce starts. No request is sent yet, because the user is still typing.
- The user types
ret. After 250ms of silence the debounce fires. - Context assembly. The prefix is everything up to the cursor, trimmed to the last 50 lines with the import block pulled back in. The suffix is the remaining 20 lines, which in this case contains the rest of the module including a
mainguard. - Signature index lookup. The index reports that
rowsis used elsewhere as a list of dicts and that a helper_cleanexists in the same file. Only the signature of_cleanis injected, not its body. - Prompt build. The three segments are joined in PSM order with ASCII pipes.
- Request.
max_tokens=48,temperature=0.1, stop at blank line, streaming on. - First token arrives at roughly 180ms and renders immediately as ghost text.
- Generation stops at the blank line after
return [...], having produced about 22 tokens. - The user presses Tab. Acceptance is recorded, and the metric feeds the nightly report.
Total wall clock from debounce fire to full render: under 400ms. Total tokens billed: roughly 900 in and 22 out. On a free tier that is comfortably sustainable for a full working day, provided the debounce and cancellation logic from section four is actually implemented. Skip the debounce and the same session generates four times the requests, most of them for prefixes the user had already typed past.
12. Related Reading
- Qwen3 Coder 480B Free API Tutorial: the large model side of repository level coding
- Why Your API Key 404s the Moment You Paste It: model names and namespace troubleshooting
- MCP Free Tutorial: wiring completion into an Agent toolchain
- Free Embedding API Complete Tutorial: the vector foundation for repository retrieval
Before you integrate, check the free API channel for the latest available code models and channels. The free API home makes filtering by category faster. If you do not have an account yet, register a free account in one click to get your keys. Treat the free API channel as the source of truth for live model status, report any dead endpoint through the free API channel as well, and bookmark the free API channel so you always see updates first, and re-check the free API channel whenever a model starts returning unexpected results.