← Back to articles
Tutorials

Free Code Completion APIs in Practice: FIM Formatting, Context Trimming, and Latency Tuning

Free Code Completion APIs in Practice: FIM Formatting, Context Trimming, and Latency Tuning

Most people who try code completion on a free tier hit the same wall on the very first attempt. They paste the whole conversation into the model and ask it to "continue." The result is a completion that lands in the wrong place, latency that climbs past two seconds, and the code sitting after the cursor getting swallowed as if it were prior context. This article explains three things properly: how to assemble the FIM prompt, how to trim the context window, and how to push latency down to something usable. Every model referenced here is one that is already verified as available on the free tier.


1. Two Different Things Both Called "Code Completion"

Inline completion and conversational coding are two distinct tasks. Using the wrong endpoint for the job is the number one cause of both latency blowups and misplaced completions.

Dimension Inline Completion Conversational Coding
Trigger Fires automatically as you type, possibly several times per second Fires when the user explicitly asks
Input shape Code before the cursor + code after the cursor A natural language question + a code snippet
Expected output A middle fill of a few to a few dozen tokens A full explanation, a whole function, multi-turn discussion
Latency budget Under 500ms, otherwise typing feels stuck 2-5 seconds is acceptable
Correct endpoint FIM / Completions Chat Completions
Typical free models mistral-code-fim-latest qwen3-coder-30b, deepseek-coder

The key conclusion: inline completion must go through the FIM format. If you drive inline completion with a Chat endpoint, the model treats your "code after the cursor" as part of the conversation history and then produces a brand new block of code that does not connect to anything. That is exactly where "the completion lands in the wrong place" comes from.


2. FIM Format: Assembling the Three Special Tokens Correctly

FIM stands for Fill-In-the-Middle. The whole idea is to split the source into three segments and use special markers to tell the model "fill the gap in the center."

2.1 Two Orderings: PSM and SPM

Ordering Structure Where it applies
PSM (Prefix-Suffix-Middle) prefix → suffix → middle The mainstream choice, default for the Mistral family
SPM (Suffix-Prefix-Middle) suffix → prefix → middle Certain StarCoder variants

The PSM assembly, using Mistral style markers as the example:

Segment Marker Content
Prefix <|fim_prefix|> Everything before the cursor
Suffix <|fim_suffix|> Everything after the cursor
Middle <|fim_middle|> Left empty; generation starts here

2.2 The Three Most Common Assembly Mistakes

  1. Do not reorder the segments. Writing prefix → middle → suffix is the single most frequent error. During training the model only ever saw prefix → suffix → middle. Get the order wrong and completion quality falls off a cliff.
  2. Do not omit the suffix. A lot of people send only the prefix, which silently degrades FIM into plain continuation. The suffix is the entire value of FIM over continuation — it tells the model what already exists downstream, so the generated variable names and indentation actually line up.
  3. The markers must be byte exact. The pipe inside <|fim_prefix|> is the ASCII pipe | (U+007C), not the full width |. Copy pasting from a Chinese language document is a reliable way to get it swapped by an input method. Once swapped, the model does not recognize the marker at all and treats it as ordinary text.

3. Context Trimming: The Life or Death Line on a Free Tier

Free tiers usually ship a context window far smaller than the paid one. Your trimming strategy decides whether the feature runs at all.

3.1 Retention Priority, Highest First

Priority Content Retention rule
1 The 50 lines before the cursor Must be kept in full
2 The 20 lines after the cursor Must be kept; this is the FIM suffix
3 The import block of the current file Pull it back in even if it sits outside those 50 lines
4 Function signatures from other files in the project Add only if budget remains
5 Historical code 50-200 lines above the cursor First thing to drop when budget is tight

3.2 Trimming by Token Budget, Not by Line Count

Never trim by line count. Trim by token count. Chinese comments and English code differ in token density by more than three times, so a line based rule will overshoot or undershoot unpredictably.

  1. Set a total budget B. For an 8K free window, reserve 1K for output, so B = 7000.
  2. Place the suffix first; call its cost s tokens.
  3. Then add the last N lines of the prefix, walking backwards from the cursor, until the running total approaches B - s.
  4. If budget remains, insert the import block at the very front of the prefix.
  5. If you are over budget, drop from the head of the prefix, never from the tail. Code closer to the cursor matters more.

3.3 Indentation Alignment: The Invisible Killer

If trimming leaves the last line of the prefix as a half statement — say you cut right through if (a && — the model will generate syntactically broken output. The rule is simple: trim boundaries must land on complete lines. Dropping one extra line is always better than keeping half of one.


4. Latency Tuning: From Two Seconds Down to 400ms

The latency budget for inline completion is brutal. People type at roughly 300ms per character, so anything above 500ms is perceived as lag.

4.1 Five Optimizations You Can Ship Today

Optimization What to do Expected gain
Cap max_tokens Set 32-64 for inline completion; never leave the default Largest single gain, often saves 60%+
Stop sequences Set stop=["\n\n", "```"] to halt at a blank line Prevents generating an entire function
Debounce Wait 250ms after typing stops before sending Cuts roughly 70% of useless requests
Request cancellation Abort the in-flight request when a new keystroke arrives Stops stale results overwriting the new position
Streaming Use stream=True and render on the first token Perceived latency drops under 200ms

4.2 Why max_tokens Is Priority Number One

Generation time scales close to linearly with output token count. Defaults are frequently 1024 or 2048, while inline completion actually needs 10 to 30 tokens. Dropping max_tokens from 1024 to 48 is a theoretical 20x difference. In practice, because time to first token is a fixed cost, end to end you typically save 60-80%.

4.3 Realistic Expectations on a Free Tier

Scenario Realistic latency band Usable for inline completion?
Free tier + small model (7B class) 300-800ms Yes, with debouncing
Free tier + large model (70B class) 1.5-4s No; conversational use only
Free tier at peak hours Fluctuates 2-5x Timeout plus silent fallback is mandatory

Conclusion: small models for inline completion, large models for conversational coding. Driving inline completion with something like qwen3-coder-30b will never meet the latency bar. That is not a configuration problem, it is physics.


5. A Runnable Example: One Complete FIM Call

This is the only place in the article that needs a code block, because the assembly order of the FIM prompt is the core teaching content.

import os
from openai import OpenAI

# Unified free tier entry point; grab the key from the console in one click
client = OpenAI(
    api_key=os.environ["APISHARE_KEY"],
    base_url="https://apishare.cc/v1",
)

def build_fim_prompt(code: str, cursor: int) -> str:
    """Split source into PSM segments at the cursor position."""
    prefix, suffix = code[:cursor], code[cursor:]
    # Trim boundaries must land on complete lines, never mid statement
    lines = prefix.split("\n")
    kept = lines[-50:] if len(lines) > 50 else lines
    prefix = "\n".join(kept)
    suffix = "\n".join(suffix.split("\n")[:20])
    return f"<|fim_prefix|>{prefix}<|fim_suffix|>{suffix}<|fim_middle|>"

def complete(code: str, cursor: int) -> str:
    resp = client.completions.create(
        model="mistral-code-fim-latest",
        prompt=build_fim_prompt(code, cursor),
        max_tokens=48,              # the make or break parameter
        temperature=0.1,            # completion wants determinism, not creativity
        stop=["\n\n", "```"],
        stream=False,
    )
    return resp.choices[0].text

if __name__ == "__main__":
    src = "def total(items):\n    s = 0\n    for it in items:\n        \n    return s\n"
    pos = src.index("        \n") + 8
    print(repr(complete(src, pos)))

Why each of those four parameters is set that way:

Parameter Value Reason
max_tokens 48 Inline completion almost never exceeds 30 tokens; 48 leaves headroom without waste
temperature 0.1 Completion needs determinism; high temperature makes variable names drift randomly
stop ["\n\n", "```"] A blank line means the logical block ended; continuing is out of bounds
model A FIM specific model Ordinary chat models were never trained on FIM markers, so the markers do nothing

6. Capability Radar: The Tradeoffs Between Three Model Classes

How to read this: there is no all round winner. The FIM specific small model dominates on latency and format support but is weak at repository level understanding. The Coder large model is the mirror image. The correct engineering answer is to wire up both — inline completion goes to the small model, while an explicit Ctrl+K question goes to the large one.


7. Troubleshooting Table

Symptom Most likely cause Diagnostic action
Completion duplicates code after the cursor Suffix was not sent; degraded into continuation Print the actual prompt and confirm there is content after <|fim_suffix|>
Completion lands in the wrong place Chat endpoint used instead of Completions Switch to the FIM endpoint and verify the model is FIM capable
Generates a huge unrelated block max_tokens left at default Drop it to 48 and set stop
Indentation is a mess Trim cut through a half statement Align trim boundaries to complete lines
Markers treated as plain text Pipe swapped to full width | by an input method Use ASCII |, copy from code rather than from docs
401 / 403 Model name pasted as a hash, or namespace prefix missing See the troubleshooting article on this site; model names must carry the namespace
Mass timeouts at peak Free tier rate limiting Add an 800ms timeout plus silent fallback; never block the editor

8. Pre Launch Acceptance Checklist

  1. Format: the three markers appear in prefix → suffix → middle order, and the pipes are ASCII
  2. Trimming: boundaries land on complete lines; prefix at most 50 lines, suffix at most 20; the import block has been pulled back in
  3. Parameters: max_tokens at most 64, temperature at most 0.2, stop configured
  4. Latency: end to end P95 under 500ms on a small model; timeout and silent fallback in place
  5. Interaction: 250ms debounce; a new keystroke aborts the previous request; first token renders via streaming
  6. Quota: daily request count tracked, with automatic switch to a backup model as the free ceiling approaches

9. Going Deeper on Repository Level Context

Everything above assumes a single file. Real projects span dozens of files, and that is where the second half of the engineering work lives.

9.1 Signature Indexing Beats Full File Injection

A common mistake is stuffing entire related files into the prompt. That burns the budget instantly and mostly adds noise. A far better approach on a constrained free tier is to index signatures only:

What to index Why it pays off
Function and method signatures The model needs names and parameter shapes, not bodies
Class declarations with base classes Reveals the inheritance contract cheaply
Exported constants and enum members Prevents the model inventing plausible but wrong names
Type aliases and interface definitions Keeps generated annotations consistent

A signature index for a medium project typically costs 500 to 1500 tokens, versus 10,000+ for the same files in full. That difference is the entire reason repository aware completion is feasible on a free tier at all.

9.2 Retrieval Order Matters More Than Retrieval Volume

When you do pull in cross file context, order it by relevance and place the most relevant chunk closest to the prefix boundary. Models weight nearby tokens more heavily. A common failure is alphabetical ordering, which puts the least relevant file right where it does the most damage.

9.3 Cache the Index, Not the Completion

The signature index changes only when files change. Cache it aggressively and invalidate on save. Completions, by contrast, should never be cached — the cursor position differs every time, and a stale completion is worse than no completion.


10. Handling Free Tier Rate Limits Gracefully

Free tiers enforce rate limits, and inline completion is the most request hungry feature in any editor. Getting this wrong means the feature works in a demo and fails in real use.

10.1 A Three Tier Degradation Ladder

Tier Condition Behavior
Normal Under quota, latency healthy Full FIM completion, streaming enabled
Throttled Approaching quota or latency above 1.2s Raise debounce to 500ms, cut max_tokens to 24
Silent off Quota exhausted or repeated 429s Stop sending requests entirely; show no error to the user

The critical design decision is the third tier. Never surface a rate limit error inside an editor. The user is typing; an error popup mid keystroke is far more disruptive than the absence of a suggestion. Degrade silently and resume automatically when the window resets.

10.2 Counting Requests Correctly

Debounce and cancellation together typically reduce request volume by 70% or more. If you are still hitting limits, the next lever is trigger policy: fire only after a word boundary or an opening bracket, rather than on every single character. That single change often halves the request count with no perceptible loss in usefulness.


11. Evaluating Completion Quality Without Guessing

"Feels good" is not a metric. Three cheap measurements will tell you whether a configuration change actually helped.

Metric Definition Target on a free tier
Acceptance rate Accepted completions divided by shown completions Above 25% is healthy
Prefix match rate Generated text that exactly continues the prefix indentation Above 90%
P95 latency 95th percentile end to end time Under 500ms for inline

Acceptance rate is the one that matters. If it sits below 15%, the problem is almost never the model — it is your trimming or your max_tokens. Fix those first, then consider switching models.



13. Language Specific Behavior You Cannot Ignore

A single set of trimming and parameter defaults will not serve every language well. Token density, comment style, and indentation semantics all differ, and those differences show up directly in completion quality.

Language family Token density Practical consequence
Python Low, whitespace significant Indentation errors are fatal; a single misaligned space breaks the block
JavaScript / TypeScript Medium Semicolon insertion and async patterns confuse models trained mostly on synchronous code
Go Medium Strict formatting means the model must emit gofmt compatible output or the diff looks noisy
Java / C# High, verbose Boilerplate dominates; the 50 line prefix window often contains only imports and class declarations
Rust High Borrow checker constraints mean a syntactically valid completion may still not compile

13.1 Adjusting the Prefix Window per Language

For verbose languages such as Java, a fixed 50 line prefix frequently contains no executable logic at all — only package declarations, imports, and annotations. The fix is to make the window semantic rather than purely positional: always include the enclosing method signature even when it sits above the 50 line cutoff. For Python, the opposite adjustment applies — the enclosing function definition plus its decorators is usually enough, and pulling in more tends to add unrelated sibling functions.

13.2 Comment Language Affects Token Cost

Comments written in Chinese consume roughly three times the tokens of equivalent English comments for the same information. On a tight free tier budget this is not a trivial detail. If your codebase has heavy Chinese comments, consider stripping comment lines from the prefix beyond the nearest 10 lines. The model rarely needs a distant comment to predict the next statement, and the token savings can be the difference between fitting the import block and not.


14. Building a Regression Test Set for Completion

Completion behavior changes when you switch models, adjust trimming, or when the provider silently updates a checkpoint. Without a test set you will not notice a regression until users complain.

14.1 What a Minimal Test Set Looks Like

Twenty cases are enough to catch most regressions. Build them from real code in your own repository rather than from synthetic examples.

Case category Count What it verifies
Function body completion 5 Basic continuation and variable naming
Argument list completion 4 Signature awareness and type correctness
Conditional branch completion 3 Control flow and indentation
Import statement completion 3 Whether the model invents nonexistent modules
Cross file reference completion 3 Whether your signature index is actually being used
Empty context edge case 2 Behavior when the prefix is trivially short

14.2 Assertions That Do Not Require an LLM Judge

Do not reach for model based scoring when deterministic checks will do. For each case assert:

  1. The output is non empty and contains no FIM marker leakage
  2. The output contains no triple backticks, which would mean it escaped the stop sequence
  3. Indentation of the first generated line matches the expected column exactly
  4. The concatenation of prefix plus generated text parses without a syntax error, checked with the language native parser
  5. End to end latency is under the configured ceiling

Assertion four is the most valuable one in the entire set. A completion that produces a syntax error when concatenated is useless regardless of how plausible it reads, and this check catches it every time at zero cost.

14.3 Running It in CI Without Burning Quota

Free tier quota is limited, so running twenty cases on every commit is not viable. Run the full set nightly, and on pull requests run only the five function body cases. Record the acceptance metrics over time; a sudden drop in the nightly run is your early warning that the provider changed something.


15. Migration Path: From Chat Based to FIM Based Completion

If you already shipped completion using a Chat endpoint, migrating to FIM is not a rewrite. The change is localized and can be done incrementally.

Step Action Risk
1 Add a feature flag that routes inline completion to the FIM path None; default stays on the old path
2 Implement build_fim_prompt and unit test the marker order in isolation Low
3 Verify the target model actually supports FIM markers with a single manual call Low
4 Enable the flag for internal users only, measure acceptance rate for three days Low
5 Compare acceptance rate and P95 latency against the Chat baseline None
6 Roll out to all users, keep the Chat path as the automatic fallback Low

Step six deserves emphasis. Keep the Chat based path alive as a fallback rather than deleting it. FIM specific models are a smaller pool than chat models, so when the FIM model is rate limited or temporarily unavailable, falling back to the chat path with a degraded but functional completion is better than showing nothing at all.

15.1 What Typically Improves, and by How Much

Based on the mechanics described above rather than on any vendor claim, the expected direction of change is consistent:

Metric Direction after migration Why
Acceptance rate Up The suffix now constrains generation to something that actually connects
Misplaced completions Down sharply The model no longer treats downstream code as history
Output length variance Down max_tokens plus stop sequences bound the response
P95 latency Down Smaller models and capped output both reduce generation time
Token cost per request Down Capped output dominates the cost calculation

The one metric that can move the wrong way is coverage. A FIM specific small model may know fewer languages than a large general chat model. That is precisely why the fallback in step six matters, and why the language breakdown in section thirteen should inform your model choice rather than being discovered afterwards.


16. A Complete Worked Session

To make the whole picture concrete, walk through one realistic editing session end to end.

The user is editing a Python utility module. They type def normalize(rows): and press enter. The editor inserts a fresh indented line and the cursor sits at column four.

  1. Debounce starts. No request is sent yet, because the user is still typing.
  2. The user types ret. After 250ms of silence the debounce fires.
  3. Context assembly. The prefix is everything up to the cursor, trimmed to the last 50 lines with the import block pulled back in. The suffix is the remaining 20 lines, which in this case contains the rest of the module including a main guard.
  4. Signature index lookup. The index reports that rows is used elsewhere as a list of dicts and that a helper _clean exists in the same file. Only the signature of _clean is injected, not its body.
  5. Prompt build. The three segments are joined in PSM order with ASCII pipes.
  6. Request. max_tokens=48, temperature=0.1, stop at blank line, streaming on.
  7. First token arrives at roughly 180ms and renders immediately as ghost text.
  8. Generation stops at the blank line after return [...], having produced about 22 tokens.
  9. The user presses Tab. Acceptance is recorded, and the metric feeds the nightly report.

Total wall clock from debounce fire to full render: under 400ms. Total tokens billed: roughly 900 in and 22 out. On a free tier that is comfortably sustainable for a full working day, provided the debounce and cancellation logic from section four is actually implemented. Skip the debounce and the same session generates four times the requests, most of them for prefixes the user had already typed past.


Before you integrate, check the free API channel for the latest available code models and channels. The free API home makes filtering by category faster. If you do not have an account yet, register a free account in one click to get your keys. Treat the free API channel as the source of truth for live model status, report any dead endpoint through the free API channel as well, and bookmark the free API channel so you always see updates first, and re-check the free API channel whenever a model starts returning unexpected results.

More in this category

Free Text Summarization API Complete Tutorial: Let LLMs Compress 1M-Word Documents into 100 WordsFree Intent Classification API Complete Tutorial: Give Your Text the Ability to Understand Human Language at Zero Cost (Verified 2026-10-07)Free Named Entity Recognition (NER) API Complete Tutorial: Extract People, Places, and Money from Text at Zero Cost (Verified 2026-10-04)Free Time Series Forecasting API Complete Tutorial: Zero-Cost “Crystal Ball” for Sales/Inventory/Energy Prices (Verified 2026-10-03)Free Semantic Textual Similarity (STS) API Complete Tutorial: Measure How Alike Two Texts Really Are at Zero Cost (Verified 2026-10-02)

Ready to use free LLM APIs?

APIShare aggregates free AI APIs worldwide — sign up and get bonus credits.