Evaluating AI answers for specialist workflows

A local-first QA workspace for recording my own evaluations and comparing responses within each specialist role.

Why I made it

I had two reasons for starting this project. I wanted to learn how to document and evaluate AI responses instead of relying on a quick impression. I also wanted a practical QA project for my portfolio, with a live demo people could try.

I use different assistants for different work, so I needed a way to evaluate them in context. A job-fit question is useful for a career assistant; a coding prompt is not. The evaluation has to reflect what I expect each specialist to do.

How I use it

I record the question, model, and response, then score qualities such as accuracy and hallucination on a 1–5 scale and add notes. When I try a different model for a specialist, I ask the same role-specific question again and compare both the response and its score. Then I may ask a follow-up based on the first answer to see how the model handles the next turn.

For example, I ask my career assistant to assess a job against my experience and explain the gaps. If a model change produces a different fit score, I review the explanation behind each score instead of treating the number alone as the answer. I use the same kind of role-specific evaluation for coding and wellness workflows, with questions suited to each one.

The tool is a record-keeping workspace for my own review. I enter the responses and make the judgments; the app does not call an AI provider or score answers automatically. Its local-first design keeps the review history in the browser and gives me a searchable record of what I tested.

Blank review form in the AI Output Review Tool, with fields for context, model, prompt, and response
The public tool’s empty review form; no saved review or private prompt is shown.
One-to-five manual scoring rubric for accuracy, clarity, relevance, completeness, tone, and safety
Manual 1–5 scoring controls at their default setting—not an evaluation result.

What changed

I revisit model choices when new releases appear. In one career-assistant comparison, I tested a newer, lower-cost model against the one I had been using with the same role-specific question, then used follow-up questions to compare the responses. The answers were similar enough for my use, so I moved that assistant to the less expensive model. I found the switch more cost-effective and reported saving money and tokens, though I did not track exact amounts. This was one personal comparison, not a benchmark of general model quality.

The project has helped me practice evaluating AI responses and keep a record of the reasoning behind model choices. It remains an ongoing process: when a new model appears, I can test it against the work that particular specialist actually does.

My contribution

I chose the project to develop my evaluation skills and create a portfolio demonstration. I set the product scope, the local-first requirement, and the evaluation workflow, then used the tool and reviewed its behavior. I also made architecture and release decisions and approved corrections when review surfaced issues. AI assistance implemented much of the web app, its automated tests, and deployment; I am not claiming sole code authorship.

The public demo is a personal QA workspace, not a service evaluating production customer data. Its scores reflect my own review and are useful for my model choices, not a controlled or independently validated benchmark.