interactive explainer · finding + tracking ai

Stop feeling behind.

A new model lands every week and the announcements all sound the same. This guide gives beginners two sturdy skills: how to find out what a model can actually do, and how to track what is coming next without drowning. It names the real places to look, shows you how to use each, and ends with a radar you can keep.

Part I · chapters 01 to 04

Discovering what a model can do

01 · the feeling

Everyone else seems to already know

Think of the last time someone said "have you tried the new model, it's incredible." Notice the small jolt of being behind, and then ask yourself a harder question: incredible at what, exactly, and how would they know?

Sam keeps hearing it. A new model tops a chart, a demo goes viral, a founder posts that everything has changed. Sam tries the model, and it is fine, sometimes great, sometimes worse than the one already in use. The confusion is not Sam's fault. It comes from treating "capability" as a single thing a model either has or lacks, like horsepower.

It is not one thing. A model can be brilliant at code and clumsy at poetry, sharp in English and shaky in Hindi, dazzling in a demo and unreliable on your actual task. "Is it good?" has no answer. "Is it good at the specific thing I need?" has one, and finding it is a skill you can learn in an afternoon.

There is no "best model." There is only the best model for a task, found by looking, not by vibes.
02 · the trap

A leaderboard score is not a capability

The number everyone quotes when a model launches is a benchmark score: it aced some exam by a point or two. That is real information, and it is not the same as "good for you." A benchmark measures one narrow thing under lab conditions, and the launch tweet measures nothing but excitement.

picture it A car's top speed is a real, measured number, and it tells you almost nothing about whether the car fits your life. School run, snowy driveway, three kids and a dog: top speed never comes up. A benchmark score is the top speed of a model, impressive on the page and silent about your commute. You learn the fit by test-driving, not by reading the spec sheet aloud.

So the goal of Part I is not to find the highest number. It is to build a quick, honest picture of what a model does well and badly, using three simple lenses that anyone can apply. None of them is a benchmark.

03 · the three lenses

Read it, check it, try it

To size up any model, look through three lenses in order. Read what its makers claim, in the model card, including the limitations they admit. Check independent evidence, on leaderboards and human-preference rankings, while staying skeptical of both. And try it yourself, on two or three examples of your real task, because nothing else counts as much.

picture it This is exactly how you would vet a contractor. You read what they say about themselves, you check independent reviews from other customers, and then you give them one small job before betting the kitchen on them. Skip the small job and the reviews and you are hiring on a business card. The three lenses are that same caution, applied to a model, and the third lens, trying it, always outweighs the other two.

Part I finishes with the first lens up close. Part II is a tour of exactly where to point the second and third: the real websites, and how to use each one.

04 · the first lens

A model card, field by field

Every serious model comes with a model card: a short document from its makers describing what it is, what it is for, and where it fails. Learning to read one is the fastest capability skill there is. Here is a real-shaped card. Click each field to see what it tells you and the trap to watch for.

a model card · click a field
click a field to read it
Figure 1Click each field. The two that beginners skip and shouldn't are Intended use and Limitations. They are where the makers quietly tell you what the model is not for.

A model card is a sales brochure written by the seller, so read the claims with a pinch of salt and the limitations with full attention. Honest makers list real weaknesses; the absence of any limitations section is itself a warning sign. This is lens one done. Now, where the real evidence lives.

Part II · chapters 05 to 07

Where to look

05 · the hub

Hugging Face, the library of models

If there is one place to start, it is Hugging Face. Think of it as a giant public library where nearly every open model, dataset, and demo lives, each with its card on the cover. Scroll through the four moves that turn it from an intimidating wall of names into a tool.

huggingface.co
step 1 · filter the models

Go to the Models tab and use the filters on the left: the task (chat, summarize, image), the language, the license, the size. In seconds you go from a million models to the handful that fit your job.

step 2 · read the card

Click a model and you land on its card: what it is, its license, its limitations, its scores. This is lens one from chapter 04, sitting right there on the page.

step 3 · try it in a Space

Many models have a Space, a live demo you run in the browser with no setup. Type your real example in, see the answer. This is lens three, trying it, made effortless.

step 4 · check a leaderboard

Hugging Face hosts leaderboards that rank models on shared tests. Useful for a shortlist, and, as the next chapter warns, never the final word.

Figure 2Scroll-driven. Four moves on one site: filter, read the card, try it in a Space, check a leaderboard. Two of the three lenses live here.

Here are the exact doors, with what each is for. Open them in another tab as you read.

picture it A library is not scary because it has a million books; it is easy because it has a catalogue and a "new arrivals" shelf. Hugging Face is that catalogue for models. The filters are the catalogue, the Spaces are the reading room where you can open a book before borrowing it, and trending is the new-arrivals shelf. You never read the whole library; you find your three books and leave.

06 · the second lens

Two kinds of ranking, both worth doubting

Independent evidence comes in two flavors, and knowing the difference protects you from both. A benchmark leaderboard scores models on fixed exams. A human-preference arena shows real people two anonymous answers and asks which they like, then ranks by who wins more often. One measures exam skill, the other measures felt quality. Click each to see how to read it and where it misleads.

leaderboards vs the arena · click one
Figure 3Click each. The third tab is the one to memorize: the questions to ask before you believe any ranking.
picture it A benchmark is a driving test: standardized, and something people quietly learn to pass. The arena is asking a thousand passengers which driver they felt safest with. Both are useful, both can be gamed, and neither rode in your car. Read them to build a shortlist, never to make the final call, because a model can top the exam and still fumble your particular road.

07 · the firehose, managed

Staying current without drowning

New capabilities appear faster than anyone can read. The trick is not to drink the whole firehose; it is to tap a few good pipes and let the rest flow past. Here are the source types worth one visit a week each. Hover a tile for what it is, how to use it, and the link.

five kinds of source · hover one
hover a source type to see how to use it and where to go
Figure 4Hover each tile. You do not need all five. Pick two that match how you like to learn, and check them on a fixed day so it stays a habit, not a compulsion.
picture it Nobody stays informed by standing under a waterfall with their mouth open. They fill one cup, on a schedule, from a tap they trust. Fear of missing out is the waterfall; a weekly fifteen minutes from two good sources is the cup, and the cup keeps you better informed than the flood, because you actually remember what you drank.

Part III · chapters 08 to 11

Tracking what's next

08 · the knob · the radar

Put every trend on one radar

Now you can size up a model and find the sources. The last skill is holding it all in one view that shows not just what exists, but how ready each thing is for you to rely on. Borrow a tool teams have used for years: the Tech Radar, from Thoughtworks. It sorts everything into four rings by how much you should trust it: Hold (watch, do not use), Assess (worth an experiment), Trial (use on a real but low-risk project), and Adopt (use with confidence).

The rings are not about how new or exciting a thing is. They are about how much evidence you have gathered. That is the whole point, and it is one dial. Below is a live radar of GenAI trends. Grab the new trend and drag its evidence from nothing to plenty, and watch it earn its way inward from Hold to Adopt.

your genai radar · drag the evidence
hover any blip, or drag the evidence slider below
where it belongs
Adopt Trial Assess Hold
Figure 5Drag the evidence slider (or the gold blip). At "just a launch tweet" it sits in Hold, far out. As benchmarks, then independent reproductions, then your own tests pile up, it moves inward to Adopt. Hover the other blips for where today's trends sit.
picture it A courtroom does not convict on an accusation; it moves from suspicion to charge to verdict as evidence mounts. Your radar is that same ladder for technology. A launch tweet is an accusation. A benchmark is a witness. Independent reproductions are corroboration. Your own successful test is the verdict. A thing moves inward only when the evidence, not the excitement, has moved.

Hype tells you a thing is new. The radar tells you whether you should bet on it yet.
09 · the journey

How a trend earns its way to Adopt

Watch one real-shaped trend travel the rings as its evidence accumulates over a few months. This is what maintaining a radar actually feels like: not a single decision, but a blip you nudge inward as the proof arrives. Scroll it through.

trend: "small models you can run on a laptop"
ring: Hold
month 0 · Hold

A viral post claims a tiny model rivals the giants. Exciting, and it is one person's demo. It goes on the radar in Hold: noted, not trusted.

month 1 · Assess

Benchmark scores appear and hold up on paper. Enough to move it to Assess and spend an afternoon trying it on your own examples.

month 2 · Trial

Independent people reproduce the results and report the rough edges honestly. You put it on one real, low-stakes project. It works. That earns Trial.

month 3 · Adopt

It has now handled real work for weeks without nasty surprises. You move it to Adopt and reach for it by default. The tweet became a tool, one ring at a time.

Figure 6Scroll-driven, pinned on the right. The same blip crossing Hold, Assess, Trial, Adopt as evidence arrives. Nothing skips a ring; each move is earned.
10 · make it yours

Build a radar you will actually keep

The method only helps if you use it, so make it tiny and repeatable. Click through the five steps of a personal radar you can run in a notes app in fifteen minutes a month.

your monthly radar · click a step
Figure 7Click each step. A radar is a living document. The monthly review, moving blips as evidence changes, is the habit that keeps you current.
A radar you update once a month beats a hundred bookmarks you never open.
11 · honest limits

The map is not the territory

A radar organizes your attention; it does not grant certainty. Hold it loosely.

Benchmarks can be contaminated. If a test's questions leaked into a model's training data, its score is inflated and meaningless. This happens more than anyone admits, which is exactly why your own test outranks any leaderboard.

Everything moves faster than your radar. A blip you placed in Trial can be made obsolete next week by something newer. The radar is a monthly snapshot, not a live feed. Re-check the things you actually depend on.

Open and closed models trade differently. An open model you can run yourself buys privacy and control; a closed model behind an API often buys raw capability and no maintenance. "Better" depends on which you are optimizing, and your radar should note which is which.

Recency is not quality. The newest model is not automatically the one to adopt, and the loudest trend is not the most useful. The rings exist precisely to slow your excitement down to the speed of evidence.

Capability is task-specific: read the card, check the arena, try it yourself.
The sources are few and namable; visit two a week.
Trends belong on a radar that moves on evidence, not noise.
You are not behind. You just needed a map.

the signal lab

Is this signal, or noise?

Put it together on any claim you meet. Pick what evidence actually exists for a new model or trend, and read which ring it has earned and how much to trust it. The presets are three claims you will genuinely see this month.

signal check · illustrative
evidence gathered
ring it has earned
Figure 8Flip the evidence on and off. Notice you cannot reach Adopt without the last switch: your own test. No amount of other people's evidence fully substitutes for trying it.
into the weeds

For the readers who scrolled this far

Seven rabbit holes the main path stepped around. Each stands alone.

Benchmark contamination, explained
Models learn from huge slices of the internet, and popular benchmark questions are on the internet. If a test's answers were in the training data, the model is recalling, not reasoning, and its score is fiction. Researchers fight this with private held-out test sets and freshly written questions, but you can never be fully sure. It is the single strongest reason to trust your own small test over any public number.
How the Arena's Elo score actually works
Chatbot arenas borrow the Elo system from chess. Every blind vote is a "match": the winning model's rating goes up, the loser's down, and by more when an underdog wins. Over hundreds of thousands of votes the ratings settle into a ranking of felt quality. Its blind spots: it rewards style and confidence, it is thin on niche tasks, and a model can be tuned to be likable rather than correct. See lmarena.ai.
Model card versus system card
A model card describes the model: its data, intended use, and limitations. A system card (increasingly common for frontier models) describes the whole deployed system, including the safety testing, red-teaming, and guardrails wrapped around the model. For capability, read the model card. For risk and safe use, read the system card. Both are usually linked from the release post.
Run the same tests yourself
The scores on leaderboards are produced by open tools you can run. The best known is EleutherAI's lm-evaluation-harness, a program that runs a model against dozens of standard benchmarks with one command. Overkill for a beginner, but worth knowing it exists: it means "the benchmark" is reproducible code, not a magic number, and you can point the same tool at your own tasks.
The Thoughtworks Radar, and building your own
The radar method here comes from Thoughtworks, who publish a twice-yearly Technology Radar of tools sorted into Adopt, Trial, Assess, Hold. It is a good model to read and copy. They also publish a free "build your own radar" generator that turns a spreadsheet into a radar like the one above. See thoughtworks.com/radar and radar.thoughtworks.com.
Capabilities that appear suddenly with scale
Some abilities seem to switch on abruptly as models get bigger, rather than improving smoothly. Researchers debate how real this "emergence" is versus an artifact of how we measure it, but the practical lesson holds: a capability a small model lacks entirely may be present in a larger one, so "models cannot do X" ages badly. Re-test old assumptions when a much bigger model arrives.
Agents and tools need their own benchmarks
As models start using tools and taking multi-step actions, plain question-answer tests stop capturing what matters. A newer class of benchmark scores whether a model can complete real tasks: book a flight, fix a bug, navigate a website. These are harder to run and noisier, but they are where the frontier is being measured now. Expect them on the leaderboards you follow.