What makes a model 'good'?
Public benchmarks can narrow the field. A small test built from your own work tells you which model is good for your business.
Every model launch says the same thing: state of the art.
You have seen the tables. Model names across the top. Tests with names like GPQA Diamond and SWE-bench down the side. Bold numbers. Shaded cells. Tiny footnotes. Every company seems to have found the arrangement where its model comes out ahead.
So what actually makes a model good?
If you only want to know what we are using, skip to the bottom. I will not be offended.
There is no single test for whether a model is good
At Runpoint, model evaluation is part of the job. Before we put a model inside a client's workflow, we need to know what it is good at, where it fails, and what kind of review it needs.
We recently worked through 103 sources for a 28-page briefing on AI benchmarks. Sam also made a five-minute explanation of how every launch can claim to be state of the art.
The short version is that there is no single event called "AI."
Public benchmarks cover a much larger range than most launch charts suggest:
- Knowledge and reasoning tests ask questions with a correct answer.
- Math and coding tests check a solution or run a test suite.
- Agent tests ask a model to browse, use tools, and complete a sequence of steps.
- Long-context tests see whether it can find and use the right information in a large pile of material.
- Professional-work tests ask for legal, medical, financial, or other domain-specific output.
- Human-preference tests compare which answer people like or trust more.
- Safety and honesty tests look for harmful behavior, deception, and other failure modes.
- Multimodal tests cover images, audio, video, and combinations of them.

The full report gets into far more of them, including specialized tests and the occasionally silly tests people invent to expose a model's quirks. Those can be useful. A weird failure sometimes teaches you more than another decimal point on a polished leaderboard. They are anecdotes, though, not a universal buying guide.


Watch: How Can Every AI Launch Be State of the Art?
Models improve fastest when there is a clear right answer
The names of the individual benchmarks matter less than the way they are graded.
Some work has an objective answer. The math is right or wrong. The code passes or fails. The browser agent reached the expected state or it did not.
This work gives a model a clean feedback signal. A lab can run a huge number of attempts, grade them immediately, and reinforce what worked. That is one reason math and coding performance can improve so quickly.
Other work needs judgment. Is this strategy wise? Is the analysis complete? Is the recommendation persuasive? Would you send this to a client? Did the model notice the thing an experienced operator would notice?
There is no universal answer key for taste, judgment, or trust. People disagree. The grader's expertise matters. The context matters. Sometimes the mistake only becomes obvious after the work has been used.
This is the central takeaway from our benchmark work: models improve fastest where the grade is clear, while much of the work businesses care about is graded by judgment.
A public benchmark can tell you that a model clears a reasonable bar. It cannot tell you whether the model will be good at your job, inside your company, with your information and your standards.
Your own work is the benchmark that matters
You do not need a research lab. You need three to five examples from a job you already understand.
Pick one recurring task. It could be summarizing a customer interview, drafting a client update, analyzing a spreadsheet, reviewing a contract, or turning meeting notes into a project plan.
Then:
- Choose three to five real examples. Remove anything sensitive.
- Give the same prompt, context, tools, and number of attempts to each model.
- Save every output. Do not rely on your memory of which one felt better.
- Review the work yourself. Ask whether it is correct, useful, and ready to use. Note how much you had to fix.
- Repeat the same test when a meaningful new model is released.
Keep the prompts, outputs, and notes in one folder. Over time, you will have a small, private benchmark built around work that actually matters to you.
You are the reviewer. That is the point.
A tiny benchmark like this will not produce an impressive global leaderboard. It will answer the question you actually have: which model should we use for this job?
Here is what we are using right now
Here is the blunt version, based on our work right now.
Something is off with Opus 5. In our use, it is confidently and frustratingly opinionated, cryptic, and just not particularly nice to work with. Your results may differ. Ours have been consistent enough that we are not reaching for it.
We rate the whole GPT-5.6 family. It has become our default. Stick with Sol (High) most of the time, but switch to others when we have simple questions or conversations. Multi-agent sessions are the key to productivity right now, since turn-by-turn has gotten a bit slower.
We love the new ChatGPT/Codex desktop app. I expected the terminal to remain the center of my workflow. I have been surprised by how quickly I moved away from it.
Sam walks through the current setup in this video. It is a longer, practical look at how the tools fit together and how we are working with agents now.

Watch Sam's current AI setup walkthrough.
The model names will change. Our process will not. Start with the public evidence, test the finalists on your own work, save the results, and repeat.
Watch Sam's five-minute benchmark explanation or read the full Runpoint benchmark briefing.