T E K D E X

Loading

TekDex is a technology-driven consulting and engineering company. We build agentic AI, cloud-native platforms and ServiceNow solutions for enterprises worldwide.

AI Engineering

Your evaluation set is the product

Teams spend months choosing a model and an afternoon deciding how they will know it works. That is exactly backwards.

TekDex11 Aug 20265 min read

Teams spend months choosing a model and an afternoon deciding how they will know it works. That is exactly backwards.

Here is a question worth asking your team: if we switched models tomorrow, how would we know whether the product got better or worse? If the honest answer is "we would try a few prompts and see how it feels", the model is not your bottleneck.

The model is the cheapest thing to change

Models are commoditising fast. What was state of the art eighteen months ago is now a cheap fallback, and the leaders change every few months. Any architecture that hard-wires you to one provider is buying a problem.

What does not commoditise is knowing what good looks like for your specific task, on your specific data, judged by your specific users. That knowledge, written down as a scored set of cases, is the durable asset. It outlives every model you will run through it.

The right model is the cheapest one that passes your evaluation set. Without the set, that sentence has no meaning.

What a usable set looks like

  • Thirty to a hundred real cases, drawn from actual usage, not invented
  • Deliberately weighted toward the awkward ones — ambiguity, missing context, questions with no good answer
  • A grading method you can defend: exact match where possible, a rubric where not, human review for the genuinely subjective
  • Version controlled, and run in CI on every prompt, model or retrieval change

The second-order effect nobody expects

Teams that build evaluation sets ship faster, and not because the model improves. They ship faster because the set removes the meeting. Nobody has to relitigate whether a change was an improvement; the number moved or it did not.

It also changes what gets built. When quality is measurable, the conversation shifts from "can AI do this?" — which is unanswerable — to "does it clear the bar on these thirty cases?", which is a Tuesday.

Choose the evaluation set with the same care you would choose the architecture. You will keep it longer.

More insights

All insights