01 / 08

A plain map of what an AI assistant can and can't do

What is AI good at, and what is it bad at?

There is one clean line between the two, and it is not the one most people guess. Once you can see it, you can predict the green lights and the red lights yourself, before you ever hit send.

"Write a toast for my sister's wedding"green
"What's the score in the match right now?"red
Read the map

01 The same machine, two different days

Same assistant, same confidence, opposite results.

Monday, you ask it to write a toast for your sister's wedding. Out comes something warm, funny, genuinely usable. You are impressed. This thing is good.

Tuesday, you ask it for the three best studies on sleep and memory, with citations. It gives you three, beautifully formatted, authors and journals and years. You look them up. Two of the papers do not exist. You are alarmed. This thing is broken.

Same assistant. Same calm, confident voice both times. So which is it, brilliant or useless? Neither. It is extremely good at one kind of task and quietly terrible at another, and the two feel identical from the outside. The whole trick is learning to tell them apart before you rely on the answer.

Quick gut checkBefore reading on, guess: is the assistant good or bad at each of these? "Rewrite my email politely." "How many days until my birthday?" "Brainstorm names for a cafe." "What is 4,829 times 7,613?" Hold your answers. We will come back to them.

02 The obvious sorting is wrong

It's not about hard versus easy.

The natural instinct is to sort tasks by difficulty. Easy things it can do, hard things it can't. That feels sensible. It is also completely wrong, and you can break it in one line.

Writing a sonnet is hard. It writes a decent one in seconds. Multiplying two four-digit numbers is something a ten-year-old with a pencil can do. On its own, it often gets that wrong. A task that stumps most humans, easy. A task any calculator nails, hard. Difficulty is clearly not the axis.

Same with "smart topics" versus "simple topics." It can explain quantum tunneling and then miscount the words in its own last sentence. The line we are looking for does not run between clever and dumb, or hard and easy. It runs somewhere else entirely, and it comes straight out of what this machine actually is.

Picture itThink of a gifted improv actor. Ask for a heartfelt speech, a fake weather report, a pirate's monologue, and it is spellbinding. Now ask, mid-scene, for the exact current time in Tokyo or the square root of 2,209. The talent that makes the improv great does nothing for those. The skill and the task have to match, and "hard" is not what decides the match.

03 The line that actually sorts everything

The only question that matters: plausible, or provably true?

Under the hood, an assistant does one thing: it guesses the words that most plausibly come next, based on the enormous pile of writing it read. That single fact draws the whole map. Scroll through the one question that sorts any task into green or red.

The task"Write a birthday poem for my mum"
The task"Give me three real studies to cite"
Would an answer that just sounds right be good enough? Or must it be verifiably true?
Green
good at it
Red
bad at it
Step 1

Start with any task at all.

Forget hard or easy. Here is something you might actually type into the box one ordinary evening.

Step 2

Ask one question about it.

Not "can it do this?" Ask instead: would an answer that merely sounds right be good enough here? Or does this have to be checkably, verifiably true?

Step 3

If sounding right is enough, green.

A birthday poem only has to land well. Warm and plausible is the entire goal, and plausible is exactly what the machine is best in the world at. Green light.

Step 4

If it must be true, red.

Three real studies have to actually exist. Here, sounding right is the trap: it will produce perfect-looking references for papers that were never written. Same machine, opposite result. Red light.

That is the whole map in one question. Not "is this hard?" but "would I be fooled by an answer that only sounds right?" When the answer is no, when plausible is genuinely good enough, you are in the green. When the answer is yes, when it has to match reality exactly, you are in the red, and the machine's greatest talent becomes the exact thing that trips you.

Green is where sounding right is enough. Red is where sounding right is a trap.

04 The live map

Pick a task. Watch the light.

Here is the map you can hold. Click any task and the signal tells you where it lands, green, amber, or red, with the reason underneath. Every reason traces back to that one question. Try to guess the light before you click, then see if the axis you just learned holds up.

The task map · click any oneexplored 0 / 21
pick a task →
The signal
waiting
Click a task on the left. The lamp lights and the reason appears here, always coming back to one thing: does this need a plausible answer, or a provably true one?

Green, it does well. Amber, it drafts well but check the risky part. Red, do not trust it without verifying elsewhere. The verdicts describe a plain assistant working on its own, before any web search or calculator is bolted on.

05 The green zone, up close

Green light: when 'plausible' is exactly what you want.

Look at what lives in the green: drafting an email, summarizing something you pasted, brainstorming twenty options, explaining a settled idea, turning notes into prose, rephrasing, translating everyday text. They look like different jobs. They are secretly the same job.

Every one of them is asking for well-shaped, plausible language, and none of them has a single hidden fact that has to be exactly right. A brainstorm wants many good-enough guesses on purpose. A summary works from words you already handed it, so there is nothing external to get wrong. An explanation of how a rainbow forms is the same in every book, so the plausible version and the correct version are the same sentence. This is not the machine getting lucky. This is the machine doing the one thing it is built to do, pointed at a task that only needs that thing.

Picture itPicture a brilliant editor and ghostwriter who has read everything, works instantly, and never gets tired, but who you would not send to the library to verify a fact. Hand that person shaping, drafting, rewording, and riffing, and they are a gift. The green zone is simply the work that plays to that person's actual strength.

It is not that these tasks are easy. It is that being roughly right and beautifully shaped is a real win for them.

06 The red zone, up close

Red light: when 'plausible' quietly becomes 'wrong.'

The red zone has four regulars, and they are worth knowing by name, because each one hands you a confident answer that is built to fool you. Flip through them. Notice that every failure is the same machine doing its one trick, aimed at a task that needed something else.

Figure · the four red-zone regularstoday's facts
You ask
What it actually does
So watch out

One pattern behind all four: it reaches for the most plausible-looking output, not a checked one. When plausible and true come apart, it cannot feel the difference.

The fourth one, its own limits, is the sneakiest. Ask "are you sure?" and it does not go back and verify. It just predicts a reasonable-sounding reply, which might be a confident "yes, I'm sure" that is still wrong, or a nervous apology that caves on a correct answer. It has no honest inner gauge of when it is reliable, so it cannot warn you. That job stays with you.

A red-zone answer is not a glitch. It is the machine succeeding at sounding right, on a task that needed being right.

07 The middle, and the escape hatch

Most real tasks are amber: draft with it, check the risky part.

Very few real requests are pure green or pure red. Plan a trip, and the shape is fantastic while the opening hours need checking. Ask for code, and it often works but must be run and tested. Draft a contract clause, and it reads like the real thing while whether it protects you is a lawyer's call. This is the amber majority: lean on it for the draft, then put your own eyes on the one part that would hurt if it were merely plausible.

There is also an escape hatch, and it is worth knowing. Many red lights turn green the moment the assistant is given a real tool. Hand it a calculator and the arithmetic becomes an actual calculation. Let it search the web and "what happened today" stops being a guess. Give it a way to run code and it can count and check instead of predict. Flip the switch below and watch the same three tasks move.

Same tasks, with and without a real toolon its own

This is why modern assistants quietly reach for a calculator, a search box, or a code runner behind the scenes. The tools exist precisely to cover the red zone.

The catchA tool helps only when the assistant actually uses one and uses it correctly. If you are not sure whether it looked something up or just guessed, treat the answer as a guess. The map does not disappear with tools, it just gets a few new green roads through the red.

08 Carrying the map with you

The one habit that keeps you safe.

You do not need to memorize which tasks are green and which are red. You need one reflex, asked before you trust any answer: would this be wrong if it just sounded right? If no, relax and enjoy the draft. If yes, that is a red or amber task, and the confident answer in front of you has earned a check, not your trust.

That single question turns a mysterious, occasionally alarming tool into a predictable one. The wedding toast, the brainstorm, the rewrite, the plain-language explainer: green, use freely. The live fact, the exact number, the real citation, the "are you sure": red, verify. Everything else: amber, draft fast, check the part that bites. The assistant did not get more or less trustworthy. You just learned where to look.

Picture itThink of it like a chainsaw versus a butter knife. Nobody calls a chainsaw "bad" because it is poor at spreading butter. You just match the tool to the cut. Knowing the green and red zones is how you stop asking the chainsaw to butter your toast, and start using it for what it is genuinely great at.

Great at sounding right. That's the gift and the whole warning.

When your task only needs a plausible, well-made answer, it is one of the most useful tools you will ever touch. When your task needs the truth, exactly, the same talent becomes a trap you now know how to spot.

Go back and answer the four gut-check tasks from the top. You can call every light now, and say why.