
How to convert daily work edits into personalized AI benchmarks
This workflow turns your regular text corrections into customized testing criteria for evaluating AI models. By rating models on your actual work requirements, you can identify which tool provides the cleanest results for your specific job.
Try it yourself
- 1Select a recurring job task and save the initial prompt alongside source files.
- 2Review the preliminary AI output and convert your standard edits into objective yes or no checklist items.
- 3Score the output manually, then ask an AI model to evaluate the same checklist without seeing your score.
- 4Refine any checklist criteria where your grade differs from the AI score, then test multiple models against the updated benchmark.
The Blend
Media startup Every reported on its initiative to replace standard industry AI tests with customized personal benchmarks for its workforce. Rather than relying on public tests that grade models on academic subjects, the organization tracks actual work tasks, such as copy editing rules or slide deck layouts, to grade model outputs against realistic employee standards.
Public benchmarks rarely show how well a software model will adapt to an individual person's specific daily workflow. By recording regular edits and converting repetitive corrections into a clear set of pass or fail criteria, workers can easily see which model handles their routine assignments best. As demonstrated by Every's team, this custom evaluation process can even reveal that smaller, less expensive models perform routine daily chores just as well as flagship systems.
However, setting up custom evaluation pipelines currently requires a structured approach and technical patience that many average office workers may struggle to maintain. It remains unclear whether tech companies will eventually build automated benchmark features directly into consumer AI applications to save users from manually logging their own edit histories. Moreover, as model vendors constantly update their software in the background, keeping personal scorecards accurate over time could easily become its own tedious chore.
Written independently by AI News Smoothie from the reporting listed below. Facts belong to the original publishers. Follow the links for their full coverage.
Ingredients
- Evals for Everyone
Creating personalized AI benchmarks allows individuals to measure model performance against their specific work preferences rather than generic academic tests.