interactive explainer · evaluating ai models

Is it actually better,
or does it just feel better?

Two chatbots give you two answers. Picking the "better" one by gut is a coin flip. This guide teaches the small, sturdy method real teams use to compare models, test prompts, and read a model's behavior, then walks you through doing every bit of it yourself in claude.ai and ChatGPT.

Part I · chapters 01 to 04

Comparing

01 · the moment

Two answers, no way to tell

Maya needs help drafting customer emails. She asks two different AI models the exact same thing and gets the two replies below. Read both, and decide, honestly, which model is better. Then notice how you decided.
same prompt: "write a polite reply declining a refund"
model A
Thank you for reaching out. After reviewing your order against our policy, we are not able to approve a refund in this case. I know that is not the answer you were hoping for, and I am happy to explain the reasoning or look at other options with you.
model B
Hi! Thanks so much for your message. Unfortunately we can't offer a refund here as it falls outside our policy window. I completely understand this is frustrating. Is there anything else I can do to help make this right?
Figure 1Two real-shaped answers to one prompt. Most people pick one in a second, then invent a reason. That speed is the problem this guide is about.

You probably had a preference within a second, then reached for a reason: this one is warmer, that one is more professional. But you judged on one example, and your reason arrived after your gut did. That is not comparison. That is a first impression wearing the costume of a decision.

Doing better is not hard, and it does not require being technical. It requires a method with three small parts, and the willingness to run more than one example. That is the whole guide.

"It feels better" is where evaluation starts, not where it ends.
02 · the trap

One example is a coin flip

Here is why a single try fools you. Models are not identical from run to run, and any one question is a fluke waiting to happen. Press the button below to run the same one-question showdown again and again, and watch the winner flip around with nothing changed.

one test case · run it repeatedly
press run
Figure 2Click run several times. Same test, and the "winner" keeps changing. If you had stopped after your first click, you would have crowned a coin toss.
picture it You are deciding which of two coffee shops is better, so you buy one drink from each on one morning. Shop A's barista was new today; Shop B nailed it. You crown Shop B. But come back ten mornings and the picture changes completely. One visit measures the morning, not the shop. One AI answer measures that roll of the dice, not the model.

03 · the method

A task, a test set, and a fair score

An evaluation, stripped to its bones, is three things. A task: the one job you actually need done, stated clearly ("decline a refund politely but firmly"). A test set: several real examples of that task, not one, so flukes cancel out. And a score: a way to judge each answer that you decide before you look, so your gut cannot rewrite the rules mid-game.

picture it This is how a cooking contest works, and for the same reason. Judges agree on the dish and the scoring (taste, presentation, follows the brief) before anyone cooks. Every chef gets the same brief. Judges taste without knowing whose plate it is. Strip any one of those away and you no longer have a contest, you have a popularity vote. A fair AI comparison is the same three rules: same task, several tries, blind scoring.

"Blind" is the quiet hero. If you know which answer came from your favorite model, you will find reasons to prefer it. Hiding the labels while you score is the cheapest way to stop fooling yourself, and you will do it by hand in Part III.

You do not need numbers-heavy rigor to start. A test set can be five examples in a document, and a score can be a thumbs up or down. The method matters more than the mathematics; running five cases beats agonizing over one.
04 · the knob · the test set

Watch the truth appear as you add tests

This is the heart of the whole method, so here it is as one slider. Below is a bench comparing two models on a task. Model A really is a little better here, but only a little, and each single test is noisy. Drag the number of test cases and watch the verdict go from a meaningless coin toss to a steady, trustworthy answer.

a/b test bench · live
Model A wins0
Model B wins0
confidence in the verdict
Figure 3Drag from 1 to 60 test cases. At a handful, the winner flip-flops and confidence stays low. Past twenty or thirty, Model A's real, small edge surfaces and stops moving. That plateau is a verdict you can trust.

Two lessons live in that slider. First, more tests do not change which model is better; they change how sure you can be. Second, the number you need is bigger than feels necessary. Five tests is far better than one, but a close call may need twenty or more before it settles. When two models are far apart, a few tests suffice; when they are close, you must look harder.

You are not measuring the model. You are measuring the model minus the noise, and noise only fades with numbers.
Part II · chapters 05 to 07

Testing prompts and reading behavior

05 · change one thing

Test prompts like a scientist, not a gambler

Comparing models is one use of the bench. The other, and the one you will use daily, is comparing prompts. When an answer disappoints, the temptation is to rewrite the whole prompt and hope. The reliable move is the opposite: keep your test set fixed, change exactly one thing, and measure whether the score moved. Click through four versions of one prompt and watch a single change at a time do the work.

one prompt, changed one step at a time
Figure 4Click through the versions. Each adds one thing and the score on the fixed test set climbs. Because only one thing changed each time, you know exactly what earned the gain.
picture it A cook whose soup is off does not change the salt, the heat, the herbs, and the pot all at once, then declare victory. They change one, taste, and learn. Change five things and the soup might improve, but you have learned nothing you can repeat. One change per test is the only way the result teaches you anything.

06 · the dice inside

Why the same prompt gives different answers

You saw the winner flip in chapter 02. Here is the reason, and it has a name and a dial. To sound natural, a model does not always pick its single most likely next word; it rolls weighted dice. How loaded those dice are is called temperature. Scroll to watch one prompt sampled again and again as the temperature rises.

prompt: "a two-word tagline for a coffee shop"
temperature: not set yet
step 0 · the prompt

One simple prompt, asked eight times. What comes back depends entirely on the temperature setting, the model's dice.

step 1 · temperature 0 · frozen

At zero, the dice are locked to the single most likely answer. Ask eight times, get the same thing eight times. Repeatable, and a little lifeless.

step 2 · temperature 0.7 · balanced

Turn it up and the model explores. Now you see a few genuinely different taglines. This is the everyday setting: varied but still sensible.

step 3 · temperature 1.3 · wild

High temperature loosens the dice all the way. Eight asks, eight different answers, some inspired and some nonsense. Creative, and a gamble.

step 4 · why it matters for testing

This is why one test lies. At any temperature above zero the model varies on its own, so to judge it fairly you must ask several times, or pin the temperature low, or both.

Figure 5Scroll-driven. The same prompt sampled eight times. As temperature rises the answers spread from identical to all-different. Identical bars mean identical outputs.
picture it Temperature is the difference between asking a careful accountant and a jazz musician to finish your sentence. The accountant gives the same sensible ending every time. The musician riffs, brilliant some nights, strange on others. Neither is "better"; they fit different jobs. You want the accountant for a policy answer and the musician for brainstorming names, and knowing which you are talking to is half of reading a model's behavior.

The everyday chat apps hide this dial and set it somewhere in the middle for you. That is a fine default, but it means the same question can give you a keeper today and a dud tomorrow. When you evaluate, that variation is the noise your test set is fighting.

07 · the tell-tale behaviors

Six behaviors every beginner should learn to spot

Reading a model means noticing how it tends to go wrong, so you can catch it. These six show up constantly. Hover each to see what it is and the quick test that reveals it.

model behaviors · hover one
hover a behavior to learn it and how to test for it
Figure 6Hover each tile. Every one comes with a one-move test you can run in any chat window in under a minute.

Notice that none of these need special tools. A made-up fact reveals itself when you ask the same question two different ways and get two different "facts." A refusal reveals itself when a rephrase gets through. You are already equipped to test all six.

Part III · chapters 08 to 11

Doing it for real

08 · the walkthrough

Run your own A/B test in ten minutes

Enough theory. Here is the exact sequence to compare two models yourself, using nothing but a browser and the free-tier chat apps. Scroll through it once, then do it for real. The panel shows what your screen looks like at each step.

your screen
step 1 · write your test set

In a plain document, write five real examples of your task. For Maya: five actual refund situations. This is your test set, and it is the part most people skip. Do not skip it.

step 2 · open the two models

Open two browser tabs. In claude.ai, pick a model from the dropdown near the message box. In ChatGPT, pick one from the model menu at the top. To compare two models on the same app, just switch the picker between runs.

step 3 · paste the same prompt

Paste test case one into both, word for word. The prompts must be identical, or you are testing your typing, not the models. Send, and keep both answers.

step 4 · regenerate to feel the noise

Under each answer, click the regenerate button two or three times. Watch the answer wobble. This is chapter 02 in your own hands, and it is why one run is not enough.

step 5 · score blind, then tally

Paste the answers into a doc with the names stripped off. Score each against your rule. Only then reveal which was which and count the wins. That number, not your first impression, is your verdict.

Figure 7Scroll-driven, pinned on the right. Five steps from blank document to a verdict you can defend. Every step uses only the free chat apps.
09 · where the buttons are

The same moves in claude.ai and ChatGPT

Both apps can do everything the method needs; they just put the controls in different places, and both hide the temperature dial on purpose. Click each move to see where it lives in each app, and where to go when you want the real lab with the dial exposed. Names and menus shift every few months, so treat this as a map, not a screenshot.

claude.ai vs chatgpt · click a move
Figure 8Click each move. Left is claude.ai, right is ChatGPT. The last row is the important one: the chat apps hide temperature; the developer consoles hand it to you.

The last move is worth the detour. When you outgrow the chat window and want to pin the temperature, set a real system prompt, and see two answers side by side, both companies offer a free developer console: the Anthropic Console for Claude and the OpenAI Playground for GPT. Same models, every dial exposed. That is where serious prompt testing happens.

10 · your homework

The five-case eval you can run today

Everything so far collapses into one small, repeatable recipe. It takes about fifteen minutes and it will change how you pick and prompt models forever.

the recipe · click a step
Figure 9Click through the five steps. This is the entire method from Part I made concrete enough to start right now.
Five real cases, scored blind, beats a thousand hot takes about which model is smartest.
11 · honest limits

A method, not a verdict machine

The bench gives you evidence, not certainty. Knowing where it falls short is what separates a useful eval from a misleading one.

Your eval is only as good as your test set. Five easy cases will crown a model that folds on the hard tenth. Choose cases that include the tricky, the edge, the ones you actually fear, not just the friendly ones.

Scores are partly subjective. "Better" for a support email is not "better" for a poem. Your scoring rule encodes your taste, and that is fine, as long as you write it down and apply it blindly and consistently.

The models move under you. Providers update models without renaming them, so a verdict from March may not hold in June. Re-run your little eval when a model changes or when the stakes are high. It is cheap.

Public benchmarks are not your task. A model topping a leaderboard was measured on someone else's test set, not yours. It is a hint, never a substitute for five cases of your own work. The weeds cover why.

A first impression is one roll of the dice.
A test set is the dice rolled enough times to see the shape.
Same task, several tries, scored blind.
That is the whole craft, and you can start with five cases today.

the confidence lab

How sure should you be?

Put the method together. Set how far apart the two models really are, how many cases you run, and whether you pinned the temperature, and read how much you can trust the verdict. The presets rebuild the three situations you will actually meet.

verdict confidence · illustrative
confidence in the verdict
effort (cases × runs)
Figure 10Try the close-call preset, then drag cases up: watch how many more you need when the models are near each other. The lazy one-shot preset shows why a single run tells you almost nothing.
into the weeds

For the readers who scrolled this far

Seven rabbit holes the main path stepped around. Each stands alone.

Temperature and top-p, precisely
Temperature scales how sharply the model favors its top guesses before it samples. Near zero it almost always takes the single most likely token, so answers repeat. Higher values flatten the odds so long-shot words get a chance. Top-p (nucleus sampling) is a cousin: it keeps only the smallest set of tokens whose probability adds up past p, then samples among those. For repeatable testing, pin temperature low. For brainstorming, raise it. The chat apps choose a middle value for you and hide the control.
Blind scoring, and why order matters too
Hiding which model wrote an answer stops brand bias. But there is a subtler trap: if Model A is always on the left, you may favor left. Shuffle the position as well as the label. Serious evals randomize both, and sometimes ask a scorer to rank rather than rate, because relative judgments ("A is better than B") are steadier than absolute ones ("A is a 7").
Using an AI to grade the AI (LLM-as-judge)
Scoring by hand does not scale past a few dozen cases, so teams often ask a strong model to score the answers against a written rubric. It is fast and surprisingly decent, but it inherits the judge's biases: it tends to prefer longer answers and its own style. Use it to triage at scale, spot-check its grades by hand, and never let the same model be both contestant and judge.
Public benchmarks: MMLU, and why the leaderboard is not your job
Benchmarks like MMLU (a big multiple-choice exam) or coding and math test sets let the field compare models on shared ground. They are useful and they are not your task. A model can ace a physics exam and still write worse refund emails than a smaller one, because your task was never on the test. Read benchmarks as a rough filter for the shortlist, then run your own five cases to actually choose.
Reasoning models change the temperature story
A newer class of models "thinks" in hidden steps before answering. For these, the useful dials and the way you test shift: you care more about whether the final answer is right than about sampling variety, and some providers restrict the temperature control entirely. The evaluation method does not change, though. Same task, several cases, blind scoring still decides it.
Reproducibility: pin the model version
"GPT" or "Claude" is a family, not a fixed thing. In the chat apps you rarely know the exact build you are talking to, and it can change overnight. If a result must be reproducible, use the developer console where you can name a specific dated model version, record it alongside your scores, and re-run against that same version later. Undated results have a quiet expiry date.
Cost and speed are part of "better"
A model that wins your quality bench by a hair but costs five times as much and answers twice as slowly may still be the wrong choice. Real evaluation scores three things at once: quality, price, and latency. The smaller, cheaper, faster model that is "good enough" on your test set is very often the right pick, and only a test set can tell you where "good enough" lands.