Two chatbots give you two answers. Picking the "better" one by gut is a coin flip. This guide teaches the small, sturdy method real teams use to compare models, test prompts, and read a model's behavior, then walks you through doing every bit of it yourself in claude.ai and ChatGPT.
You probably had a preference within a second, then reached for a reason: this one is warmer, that one is more professional. But you judged on one example, and your reason arrived after your gut did. That is not comparison. That is a first impression wearing the costume of a decision.
Doing better is not hard, and it does not require being technical. It requires a method with three small parts, and the willingness to run more than one example. That is the whole guide.
Here is why a single try fools you. Models are not identical from run to run, and any one question is a fluke waiting to happen. Press the button below to run the same one-question showdown again and again, and watch the winner flip around with nothing changed.
An evaluation, stripped to its bones, is three things. A task: the one job you actually need done, stated clearly ("decline a refund politely but firmly"). A test set: several real examples of that task, not one, so flukes cancel out. And a score: a way to judge each answer that you decide before you look, so your gut cannot rewrite the rules mid-game.
"Blind" is the quiet hero. If you know which answer came from your favorite model, you will find reasons to prefer it. Hiding the labels while you score is the cheapest way to stop fooling yourself, and you will do it by hand in Part III.
This is the heart of the whole method, so here it is as one slider. Below is a bench comparing two models on a task. Model A really is a little better here, but only a little, and each single test is noisy. Drag the number of test cases and watch the verdict go from a meaningless coin toss to a steady, trustworthy answer.
Two lessons live in that slider. First, more tests do not change which model is better; they change how sure you can be. Second, the number you need is bigger than feels necessary. Five tests is far better than one, but a close call may need twenty or more before it settles. When two models are far apart, a few tests suffice; when they are close, you must look harder.
Comparing models is one use of the bench. The other, and the one you will use daily, is comparing prompts. When an answer disappoints, the temptation is to rewrite the whole prompt and hope. The reliable move is the opposite: keep your test set fixed, change exactly one thing, and measure whether the score moved. Click through four versions of one prompt and watch a single change at a time do the work.
You saw the winner flip in chapter 02. Here is the reason, and it has a name and a dial. To sound natural, a model does not always pick its single most likely next word; it rolls weighted dice. How loaded those dice are is called temperature. Scroll to watch one prompt sampled again and again as the temperature rises.
One simple prompt, asked eight times. What comes back depends entirely on the temperature setting, the model's dice.
At zero, the dice are locked to the single most likely answer. Ask eight times, get the same thing eight times. Repeatable, and a little lifeless.
Turn it up and the model explores. Now you see a few genuinely different taglines. This is the everyday setting: varied but still sensible.
High temperature loosens the dice all the way. Eight asks, eight different answers, some inspired and some nonsense. Creative, and a gamble.
This is why one test lies. At any temperature above zero the model varies on its own, so to judge it fairly you must ask several times, or pin the temperature low, or both.
The everyday chat apps hide this dial and set it somewhere in the middle for you. That is a fine default, but it means the same question can give you a keeper today and a dud tomorrow. When you evaluate, that variation is the noise your test set is fighting.
Reading a model means noticing how it tends to go wrong, so you can catch it. These six show up constantly. Hover each to see what it is and the quick test that reveals it.
Notice that none of these need special tools. A made-up fact reveals itself when you ask the same question two different ways and get two different "facts." A refusal reveals itself when a rephrase gets through. You are already equipped to test all six.
Enough theory. Here is the exact sequence to compare two models yourself, using nothing but a browser and the free-tier chat apps. Scroll through it once, then do it for real. The panel shows what your screen looks like at each step.
In a plain document, write five real examples of your task. For Maya: five actual refund situations. This is your test set, and it is the part most people skip. Do not skip it.
Open two browser tabs. In claude.ai, pick a model from the dropdown near the message box. In ChatGPT, pick one from the model menu at the top. To compare two models on the same app, just switch the picker between runs.
Paste test case one into both, word for word. The prompts must be identical, or you are testing your typing, not the models. Send, and keep both answers.
Under each answer, click the regenerate button two or three times. Watch the answer wobble. This is chapter 02 in your own hands, and it is why one run is not enough.
Paste the answers into a doc with the names stripped off. Score each against your rule. Only then reveal which was which and count the wins. That number, not your first impression, is your verdict.
Both apps can do everything the method needs; they just put the controls in different places, and both hide the temperature dial on purpose. Click each move to see where it lives in each app, and where to go when you want the real lab with the dial exposed. Names and menus shift every few months, so treat this as a map, not a screenshot.
The last move is worth the detour. When you outgrow the chat window and want to pin the temperature, set a real system prompt, and see two answers side by side, both companies offer a free developer console: the Anthropic Console for Claude and the OpenAI Playground for GPT. Same models, every dial exposed. That is where serious prompt testing happens.
Everything so far collapses into one small, repeatable recipe. It takes about fifteen minutes and it will change how you pick and prompt models forever.
The bench gives you evidence, not certainty. Knowing where it falls short is what separates a useful eval from a misleading one.
Your eval is only as good as your test set. Five easy cases will crown a model that folds on the hard tenth. Choose cases that include the tricky, the edge, the ones you actually fear, not just the friendly ones.
Scores are partly subjective. "Better" for a support email is not "better" for a poem. Your scoring rule encodes your taste, and that is fine, as long as you write it down and apply it blindly and consistently.
The models move under you. Providers update models without renaming them, so a verdict from March may not hold in June. Re-run your little eval when a model changes or when the stakes are high. It is cheap.
Public benchmarks are not your task. A model topping a leaderboard was measured on someone else's test set, not yours. It is a hint, never a substitute for five cases of your own work. The weeds cover why.
A first impression is one roll of the dice.
A test set is the dice rolled enough times to see the shape.
Same task, several tries, scored blind.
That is the whole craft, and you can start with five cases today.
Put the method together. Set how far apart the two models really are, how many cases you run, and whether you pinned the temperature, and read how much you can trust the verdict. The presets rebuild the three situations you will actually meet.
Seven rabbit holes the main path stepped around. Each stands alone.