A plain map of what an AI assistant can and can't do
There is one clean line between the two, and it is not the one most people guess. Once you can see it, you can predict the green lights and the red lights yourself, before you ever hit send.
01 The same machine, two different days
Monday, you ask it to write a toast for your sister's wedding. Out comes something warm, funny, genuinely usable. You are impressed. This thing is good.
Tuesday, you ask it for the three best studies on sleep and memory, with citations. It gives you three, beautifully formatted, authors and journals and years. You look them up. Two of the papers do not exist. You are alarmed. This thing is broken.
Same assistant. Same calm, confident voice both times. So which is it, brilliant or useless? Neither. It is extremely good at one kind of task and quietly terrible at another, and the two feel identical from the outside. The whole trick is learning to tell them apart before you rely on the answer.
02 The obvious sorting is wrong
The natural instinct is to sort tasks by difficulty. Easy things it can do, hard things it can't. That feels sensible. It is also completely wrong, and you can break it in one line.
Writing a sonnet is hard. It writes a decent one in seconds. Multiplying two four-digit numbers is something a ten-year-old with a pencil can do. On its own, it often gets that wrong. A task that stumps most humans, easy. A task any calculator nails, hard. Difficulty is clearly not the axis.
Same with "smart topics" versus "simple topics." It can explain quantum tunneling and then miscount the words in its own last sentence. The line we are looking for does not run between clever and dumb, or hard and easy. It runs somewhere else entirely, and it comes straight out of what this machine actually is.
03 The line that actually sorts everything
Under the hood, an assistant does one thing: it guesses the words that most plausibly come next, based on the enormous pile of writing it read. That single fact draws the whole map. Scroll through the one question that sorts any task into green or red.
Forget hard or easy. Here is something you might actually type into the box one ordinary evening.
Not "can it do this?" Ask instead: would an answer that merely sounds right be good enough here? Or does this have to be checkably, verifiably true?
A birthday poem only has to land well. Warm and plausible is the entire goal, and plausible is exactly what the machine is best in the world at. Green light.
Three real studies have to actually exist. Here, sounding right is the trap: it will produce perfect-looking references for papers that were never written. Same machine, opposite result. Red light.
That is the whole map in one question. Not "is this hard?" but "would I be fooled by an answer that only sounds right?" When the answer is no, when plausible is genuinely good enough, you are in the green. When the answer is yes, when it has to match reality exactly, you are in the red, and the machine's greatest talent becomes the exact thing that trips you.
Green is where sounding right is enough. Red is where sounding right is a trap.
04 The live map
Here is the map you can hold. Click any task and the signal tells you where it lands, green, amber, or red, with the reason underneath. Every reason traces back to that one question. Try to guess the light before you click, then see if the axis you just learned holds up.
Green, it does well. Amber, it drafts well but check the risky part. Red, do not trust it without verifying elsewhere. The verdicts describe a plain assistant working on its own, before any web search or calculator is bolted on.
05 The green zone, up close
Look at what lives in the green: drafting an email, summarizing something you pasted, brainstorming twenty options, explaining a settled idea, turning notes into prose, rephrasing, translating everyday text. They look like different jobs. They are secretly the same job.
Every one of them is asking for well-shaped, plausible language, and none of them has a single hidden fact that has to be exactly right. A brainstorm wants many good-enough guesses on purpose. A summary works from words you already handed it, so there is nothing external to get wrong. An explanation of how a rainbow forms is the same in every book, so the plausible version and the correct version are the same sentence. This is not the machine getting lucky. This is the machine doing the one thing it is built to do, pointed at a task that only needs that thing.
It is not that these tasks are easy. It is that being roughly right and beautifully shaped is a real win for them.
06 The red zone, up close
The red zone has four regulars, and they are worth knowing by name, because each one hands you a confident answer that is built to fool you. Flip through them. Notice that every failure is the same machine doing its one trick, aimed at a task that needed something else.
One pattern behind all four: it reaches for the most plausible-looking output, not a checked one. When plausible and true come apart, it cannot feel the difference.
The fourth one, its own limits, is the sneakiest. Ask "are you sure?" and it does not go back and verify. It just predicts a reasonable-sounding reply, which might be a confident "yes, I'm sure" that is still wrong, or a nervous apology that caves on a correct answer. It has no honest inner gauge of when it is reliable, so it cannot warn you. That job stays with you.
A red-zone answer is not a glitch. It is the machine succeeding at sounding right, on a task that needed being right.
07 The middle, and the escape hatch
Very few real requests are pure green or pure red. Plan a trip, and the shape is fantastic while the opening hours need checking. Ask for code, and it often works but must be run and tested. Draft a contract clause, and it reads like the real thing while whether it protects you is a lawyer's call. This is the amber majority: lean on it for the draft, then put your own eyes on the one part that would hurt if it were merely plausible.
There is also an escape hatch, and it is worth knowing. Many red lights turn green the moment the assistant is given a real tool. Hand it a calculator and the arithmetic becomes an actual calculation. Let it search the web and "what happened today" stops being a guess. Give it a way to run code and it can count and check instead of predict. Flip the switch below and watch the same three tasks move.
This is why modern assistants quietly reach for a calculator, a search box, or a code runner behind the scenes. The tools exist precisely to cover the red zone.
08 Carrying the map with you
You do not need to memorize which tasks are green and which are red. You need one reflex, asked before you trust any answer: would this be wrong if it just sounded right? If no, relax and enjoy the draft. If yes, that is a red or amber task, and the confident answer in front of you has earned a check, not your trust.
That single question turns a mysterious, occasionally alarming tool into a predictable one. The wedding toast, the brainstorm, the rewrite, the plain-language explainer: green, use freely. The live fact, the exact number, the real citation, the "are you sure": red, verify. Everything else: amber, draft fast, check the part that bites. The assistant did not get more or less trustworthy. You just learned where to look.
When your task only needs a plausible, well-made answer, it is one of the most useful tools you will ever touch. When your task needs the truth, exactly, the same talent becomes a trap you now know how to spot.
Go back and answer the four gut-check tasks from the top. You can call every light now, and say why.