Comparing models and revising prompts.
Make AI quality something that can be examined and improved.
Where this work fits
A production model was being deprecated, and its replacement threatened a threefold cost increase. I needed to compare alternatives without trading away quality: fluent answers could still invent a qualification, change a job title, or omit useful experience.
I compared model outputs using the same source examples and explicit quality criteria. I inspected unsupported claims and missing experience behind the scores, combined model judging with evidence checks, and reran comparisons after prompt revisions to look for improvements and regressions.
Compare answers against evidence
- 01Same source
- 02Two answers
- 03Cross-judge
- 04Revise prompt
- 05Score again
Built inventory & invoicing APIs with Node.js.
Job asks for Kubernetes. The candidate’s experience does not include it.I build inventory & invoicing APIs with Node.js.
I build inventory & invoicing APIs with Node.js.
The revised answers restore Node.js and the missing inventory and invoicing work. Both criterion scores move from zero to two. Scores still need an evidence check.
How the scoring works
Each answer is checked for two things: whether its claims are supported, and whether it preserves relevant experience. A score of 0 fails the check, 1 partially meets it, and 2 meets it.
Both answers are checked against the same profile. One adds a skill the candidate never claimed; the other drops relevant experience. Revising the prompt addresses both mistakes, then the answers are scored again.
A higher score was not enough.
The evaluations exposed a tension between preventing invention and preserving useful source material. A stricter instruction could remove content it should have kept. Cross-model judging provided another perspective, but I still checked the evidence and used deterministic validation for rules with a definite answer.
The skills behind the work
- LLM evaluation
- Compare outputs against repeatable examples and explicit quality criteria.
- Grounding
- Check that generated claims are supported by the source information.
- Prompt engineering
- Refine model instructions using observed behavior and failure cases.
I shipped the highest-performing integration from the comparison. It delivered 43% higher evaluation quality at 23% lower model cost. The process also exposed prompt failures that I could correct and test again.
- Evaluation quality
- +43%Quality improvement from the selected model integration after comparing three LLM providers.
- Model cost
- −23%Lower cost for the selected integration in the model comparison.