Match model size to the job — grounded retrieval and phrasing does not need a frontier model.
A strong, pre-grounded system prompt beats a bigger model for a narrow domain.
For a free public feature, first-token latency and cost per message matter as much as answer quality.
The job the model actually does
The twin answers questions like “what's your experience?” or “what have you built?” The answers already exist in my CMS. So the model isn't reasoning — it's selecting the relevant grounded facts and phrasing them in my voice. That's a task an 8B model does well.
Why not the bigger model
70b writes nicer open-ended prose, but it costs more per message and its first token is slower. For a free public feature the latency and cost matter, and the extra reasoning headroom is capacity the task barely touches. The right lever wasn't a bigger model — it was a better system prompt.