← The Torre work01 / TORRE · CONTRIBUTION

Comparing models and revising prompts.

Make AI quality something that can be examined and improved.

Where this work fits

A production model was being deprecated, and its replacement threatened a threefold cost increase. I needed to compare alternatives without trading away quality: fluent answers could still invent a qualification, change a job title, or omit useful experience.

MY CONTRIBUTION

I compared model outputs using the same source examples and explicit quality criteria. I inspected unsupported claims and missing experience behind the scores, combined model judging with evidence checks, and reran comparisons after prompt revisions to look for improvements and regressions.

THE IDEA, ILLUSTRATED

Compare answers against evidence

  1. 01Same source
  2. 02Two answers
  3. 03Cross-judge
  4. 04Revise prompt
  5. 05Score again
CANDIDATE PROFILE · SOURCE OF TRUTH

Built inventory & invoicing APIs with Node.js.

Job asks for Kubernetes. The candidate’s experience does not include it.
Model AREVISED ANSWERModel BJUDGES

I build inventory & invoicing APIs with Node.js.

02/ 2Grounded claims
Unsupported skill removed. Source fact restored.
Model BREVISED ANSWERModel AJUDGES

I build inventory & invoicing APIs with Node.js.

02/ 2Preserved experience
The candidate’s actual work is visible again.
Neither model grades itself.

The revised answers restore Node.js and the missing inventory and invoicing work. Both criterion scores move from zero to two. Scores still need an evidence check.

How the scoring works

Each answer is checked for two things: whether its claims are supported, and whether it preserves relevant experience. A score of 0 fails the check, 1 partially meets it, and 2 meets it.

Both answers are checked against the same profile. One adds a skill the candidate never claimed; the other drops relevant experience. Revising the prompt addresses both mistakes, then the answers are scored again.

THE ENGINEERING DECISION

A higher score was not enough.

The evaluations exposed a tension between preventing invention and preserving useful source material. A stricter instruction could remove content it should have kept. Cross-model judging provided another perspective, but I still checked the evidence and used deterministic validation for rules with a definite answer.

The skills behind the work

LLM evaluation
Compare outputs against repeatable examples and explicit quality criteria.
Grounding
Check that generated claims are supported by the source information.
Prompt engineering
Refine model instructions using observed behavior and failure cases.
THE WIDER RESULT

I shipped the highest-performing integration from the comparison. It delivered 43% higher evaluation quality at 23% lower model cost. The process also exposed prompt failures that I could correct and test again.

Evaluation quality
+43%Quality improvement from the selected model integration after comparing three LLM providers.
Model cost
−23%Lower cost for the selected integration in the model comparison.
READ THIS IN CONTEXTI tested what a better-looking answer might hide. →
Related contributionsAsk questions that actually change the decisionI contributed to AI-assisted screening generation and the experience of reviewing and editing questions. Structured output, validation, and fallback handling helped turn generated text into questions the product could reliably use.