What is Fine-Tuning?
Agent InternalsAdapting a model on your own examples so its behaviour changes: the format it returns, the vocabulary it uses, the tone it holds. It changes how a model responds rather than what it knows, which is why it is not a way to add facts.
Why It Matters
Fine-tuning is the answer teams reach for when a model “does not understand our domain”. Usually the actual problem is that it has not been given the material, which is a retrieval problem, and fine-tuning it instead produces an expensive model that is fluent in the wrong facts.
The distinction is worth holding onto: retrieval supplies knowledge, fine-tuning shapes behaviour. A support assistant needs the current returns policy, which changes and therefore belongs in retrieval. It may also need to answer in a house style and a fixed structure, which is what fine-tuning is good at.
What It Changes
Format. Returning a consistent structure rather than a plausible paragraph, which removes a parsing layer from everything downstream.
Vocabulary. Using your terms, your product names, and your abbreviations without being told each time.
Tone and refusal behaviour. Holding a register, and declining in the way your organisation would.
Prompt length. A model that already knows the house style needs less instruction per call, which reduces cost and latency on every request.
What it does not do is reliably add facts. A model trained on a document can reproduce its phrasing and still get a detail wrong, with no source to check against, which is the worst of both worlds in a regulated workflow.
How to Decide
Work up the ladder rather than starting at the top. Prompting first, because a clearer instruction solves more cases than teams expect. Retrieval next, because most “the model does not know our business” complaints are a missing document. Fine-tuning when the requirement is consistent behaviour at volume, and the evaluation set exists to prove the change is an improvement.
The honest signals for it: a fixed output format that a prompt cannot hold reliably, a domain vocabulary dense enough that every prompt is mostly glossary, or a call volume where the shorter prompt pays for the training.
Where It Breaks
Fine-tuning to add knowledge. The most expensive mistake in this area. The model cannot cite what it learned, cannot be corrected without retraining, and can still produce a wrong detail in a confident register.
No evaluation set. A fine-tuned model is a new model, and a change nobody measured is a change nobody can defend. The golden dataset and the regression gate matter more here than anywhere else.
No versioning. Once training runs, the model is an artefact. If it is not versioned alongside the data that produced it, a behaviour change cannot be rolled back and a cause cannot be traced.
Training data nobody reviewed. The examples become the specification. If they contain a shortcut, an inconsistency, or personal data, the model learns it, and the provenance of that data is what an auditor will ask for.
Drift against a moving source. A model tuned on last year’s policies holds them. Without a retraining trigger tied to the source of truth, it becomes confidently stale.
How Flytebit Handles It
We start with prompting and retrieval, and reach for fine-tuning when the requirement is behavioural and the measurement exists to support it. When we do train, the model is treated as a versioned artefact with the training data recorded, an eval gate before release, and a rollback path to the previous version. The retrieval alternative is covered in our RAG development work, and the wider build in generative AI development.