← The Torre work01 / TORRE · CONTRIBUTION

Ask questions that actually change the decision.

Make screening questions clearer and more useful.

Where this work fits

A production model was being deprecated, and its replacement threatened a threefold cost increase. I needed to compare alternatives without trading away quality: fluent answers could still invent a qualification, change a job title, or omit useful experience.

MY CONTRIBUTION

I contributed to AI-assisted screening generation and the experience of reviewing and editing questions. Structured output, validation, and fallback handling helped turn generated text into questions the product could reliably use.

THE IDEA, ILLUSTRATED

Compare answers against evidence

  1. 01Same source
  2. 02Two answers
  3. 03Cross-judge
  4. 04Revise prompt
  5. 05Score again
CANDIDATE PROFILE · SOURCE OF TRUTH

Built inventory & invoicing APIs with Node.js.

Job asks for Kubernetes. The candidate’s experience does not include it.
Model AREVISED ANSWERModel BJUDGES

I build inventory & invoicing APIs with Node.js.

02/ 2Grounded claims
Unsupported skill removed. Source fact restored.
Model BREVISED ANSWERModel AJUDGES

I build inventory & invoicing APIs with Node.js.

02/ 2Preserved experience
The candidate’s actual work is visible again.
Neither model grades itself.

The revised answers restore Node.js and the missing inventory and invoicing work. Both criterion scores move from zero to two. Scores still need an evidence check.

How the scoring works

Each answer is checked for two things: whether its claims are supported, and whether it preserves relevant experience. A score of 0 fails the check, 1 partially meets it, and 2 meets it.

THE ENGINEERING DECISION

A higher score was not enough.

The evaluations exposed a tension between preventing invention and preserving useful source material. A stricter instruction could remove content it should have kept. Cross-model judging provided another perspective, but I still checked the evidence and used deterministic validation for rules with a definite answer.

The skills behind the work

Structured LLM output
Give generated content a predictable shape the application can validate and use.
Prompt engineering
Refine model instructions using observed behavior and failure cases.
Validation
Reject invalid or unsupported input before it affects the rest of the application.
THE WIDER RESULT

I shipped the highest-performing integration from the comparison. It delivered 43% higher evaluation quality at 23% lower model cost. The process also exposed prompt failures that I could correct and test again.

READ THIS IN CONTEXTI tested what a better-looking answer might hide. →
Related contributionsComparing models and revising promptsI compared model outputs using the same source examples and explicit quality criteria. I inspected unsupported claims and missing experience behind the scores, combined model judging with evidence checks, and reran comparisons after prompt revisions to look for improvements and regressions.