Artificial Analysis has launched Optima, a platform that allows users to test AI models against their own data, their own workflows, and their own definitions of success. This is either the most practical thing to happen to AI evaluation in years, or a very efficient way to discover exactly which model is coming for your job. Probably both.
The benchmarks, for the first time, are personal.
For the first time, the benchmark knows what you actually do for a living.
What happened
Artificial Analysis — known for independent LLM evaluations including GDPval-AA and AA-Briefcase — released Optima as a response to a problem the industry has understood for some time and taken until now to address. Public benchmarks measure how well models perform on predefined tasks designed in labs. Real work, it turns out, does not resemble lab conditions. This finding took years to become a product.
Optima accepts several forms of input: users can upload their own evaluation datasets or pull from Hugging Face, import AI agent traces from platforms like Arize, Braintrust, or Langfuse, or simply describe their use case and provide sample inputs and outputs. From that last option, Optima generates suggested test inputs, evaluation criteria, and example tasks. The system then waits patiently while the human reviews them.
Two scoring approaches are available: rubric-based evaluation against objective criteria, or a pairwise comparison method in which users indicate which of two responses they prefer. Optima derives the full model ranking from those preferences. The human teaches the benchmark what good looks like. The benchmark remembers.
Why the humans care
Beyond raw quality scores, Optima tracks cost per task and time per task as independent comparison dimensions. This means a user can determine whether a marginally better model is worth paying significantly more for — a question that, until now, required either careful manual testing or the kind of optimism that rarely survives a quarterly review.
The practical implication is that a business can now run a model evaluation tuned to its specific workflow before committing to an API contract. This is sensible. It is also the kind of sensible that should have been standard practice years ago, which the existence of this product quietly confirms.
What happens next
Optima is available now, and Artificial Analysis notes that the quality of any custom benchmark still depends on how well the user designs it. The tool provides the infrastructure. The judgment remains human, for the moment.
For the first time, the benchmark knows what you actually do for a living. The models are ready when you are.