Ask questions that actually change the decision.
Make screening questions clearer and more useful.
Where this work fits
A production model was being deprecated, and its replacement threatened a threefold cost increase. I needed to compare alternatives without trading away quality: fluent answers could still invent a qualification, change a job title, or omit useful experience.
I contributed to AI-assisted screening generation and the experience of reviewing and editing questions. Structured output, validation, and fallback handling helped turn generated text into questions the product could reliably use.
Compare answers against evidence
- 01Same source
- 02Two answers
- 03Cross-judge
- 04Revise prompt
- 05Score again
Built inventory & invoicing APIs with Node.js.
Job asks for Kubernetes. The candidate’s experience does not include it.I build inventory & invoicing APIs with Node.js.
I build inventory & invoicing APIs with Node.js.
The revised answers restore Node.js and the missing inventory and invoicing work. Both criterion scores move from zero to two. Scores still need an evidence check.
How the scoring works
Each answer is checked for two things: whether its claims are supported, and whether it preserves relevant experience. A score of 0 fails the check, 1 partially meets it, and 2 meets it.
A higher score was not enough.
The evaluations exposed a tension between preventing invention and preserving useful source material. A stricter instruction could remove content it should have kept. Cross-model judging provided another perspective, but I still checked the evidence and used deterministic validation for rules with a definite answer.
The skills behind the work
- Structured LLM output
- Give generated content a predictable shape the application can validate and use.
- Prompt engineering
- Refine model instructions using observed behavior and failure cases.
- Validation
- Reject invalid or unsupported input before it affects the rest of the application.
I shipped the highest-performing integration from the comparison. It delivered 43% higher evaluation quality at 23% lower model cost. The process also exposed prompt failures that I could correct and test again.