How AI Works

Which AI model should you choose? Wrong question — here's the method that always works

Every model update improves something and breaks something else. My observations on model "degradation" and a task-based selection method that outlives any release.

Which AI model should you choose? Wrong question — here's the method that always works

01 — The observation

Why can a new model be worse than the old one?

With every new release and every tweak to how a model talks to users, it "degrades" at some other set of tasks. That's not paranoia — it's the price of any improvement.

From my daily practice: after ChatGPT got memory, it pulled my data and projects without a thousand explanations — but because of memory it sometimes becomes a repeating parrot: the same things, in the same words. It got an answer-style training layer — and started writing short, punchy, clear. But that reduced stylistic range: ask for different rhythms, get similar texts.

Developers always trade something: accuracy vs creativity, obedience vs character, brevity vs depth. So "which model is best" has no answer — while "which is best for this task today" does, and the answer changes with every release.

02 — Examples

What did this look like on real models?

A few examples from my tasks — historical by now (models have changed several times since), but the pattern is timeless:

The "improved ↔ broke" pattern on live versions
Claude Sonnet 3.5–3.7 → magnificent at texts, metaphors,
  emotion; made logic mistakes
Claude Sonnet 4       → fewer logic mistakes;
  lost depth and stylistic range

Gemini (early)        → weak at texts, strong at analysis
Gemini Flash 2.5      → free, yet beat paid models
  at text style in many tasks

ChatGPT 4 (old)       → sometimes wrote better than newer versions

The takeaway: "boosted logic — depth sagged," "tuned the style — a free model beat the paid ones." A model's strengths aren't written on the price tag and don't track the release date. They're only visible on your task.

03 — The method

What do you do when a proven prompt starts producing nonsense?

My conclusion is singular. When I'm 1000% sure about my prompt — it delivered perfect results for ages — and now it outputs garbage, I don't sit rewriting the prompt endlessly.

I simply go and test a bunch of other models on the same task. To find out: what am I using for this TODAY?

The usual way

Prompt broke → tweak wording in the same model for hours → get angry → conclude "AI got dumber."

The faster way

Prompt broke → same prompt into 3–4 other models → compare → work where the result is better. 15 minutes instead of an evening.

There are comparison services that send one request to several models and show the answers side by side — I've used promptcannon.com. But even without them, the test takes a quarter of an hour: same inputs, by hand, into 3–4 chats.

04 — The guide

How do you run a model test on your own task?

Walk the test once — and the method stays with you forever:

Not an abstract "write a post," but your live task where you know what good looks like: a post in your voice, a form review, a deck structure. The reference is your measuring instrument.

The same prompt, the same materials, the same order — into every model, unchanged. If you adapt the wording per model, the comparison collapses: you'd be comparing prompts, not models.

The minimal set: ChatGPT, Claude, Gemini — plus whatever you use now. By hand in four tabs or via a comparison service. Put the answers side by side in one document.

Not "which answer is prettier," but 3–4 criteria of your task: voice match, factual accuracy, structure, edits-to-publishable. The winner isn't the "smartest" — it's the one whose output you'd use with minimal edits.

"Texts — X, analysis — Y, structures — Z. Verified: September." The date matters: the next major release can flip the map, and then the test repeats — 15 minutes, not a research project.

05 — The map

What does a personal model map look like?

After a couple of tests you get your own map: which model for which task type. Mine looks roughly like this (yours will differ — that's the point):

Texts in my voice        → model A (depth, metaphors)
Logic, structures, code  → model B (fewer mistakes)
Fast drafts, volume      → model C (cheap and quick)
Data analysis, tables    → model B
Ideas and brainstorming  → models A + C in parallel

Two rules for the map. First: it lives with a date — releases redraw it a few times a year. Second: don't move "your whole life" to one model because it's trendy. Move per task — to wherever that task is solved best.

  • Listed the 3–5 task types I do with AI weekly
  • Tested each type across 3–4 models with identical inputs
  • Wrote down "task → model" verdicts with a verification date
  • Set a reminder to re-check the map after a major release
  • Proven prompt broke → model test first, prompt surgery second

06 — The stance

Why does loyalty to one model cost you?

Most people pick a model once — by habit or by subscription — and then quietly suffer: "it writes worse somehow." Meanwhile the model genuinely changed; nobody just checked.

Switching tools isn't betrayal or chaos. It's professional hygiene: a photographer owns several lenses, a chef several knives. A subscription doesn't oblige you to do everything in one model — especially when a strong free model can beat paid ones on your task, as happened for me with text style.

Take this — the rule in one line
The task picks the tool, not the subscription.
Prompt broke → a 15-minute test across 3–4 models →
work where the result is better TODAY.
Takeaway

There is no best model — only the best one for your task today. Every release improves something and breaks something, so the method is one: proven prompt outputs nonsense → a 15-minute test across 3–4 models → work where the result is better. Build your model map with the checklist — and refresh it after major releases.

FAQ

Which AI model is currently the best?

None — the question is framed wrong. Each model is strong at something: one goes deeper in texts, another is more precise in logic, a third is better at analysis. And the layout shifts with every release: updates improve one thing and break another. Only a test on your specific task gives a working answer.

Why did the model "get dumber" when my prompt never changed?

Because the model was updated: communication style tuned, memory added, fine-tuning changed. Gains in one area almost always cost degradation in another — e.g. texts get shorter and clearer but lose stylistic range. The prompt isn't guilty — the performer changed.

How do I quickly test several models?

The same prompt with the same materials into 3–4 models, answers side by side, compared on your criteria (voice, accuracy, edits needed). By hand it's 15 minutes; comparison services can send one request to several models at once. The key: don't change the inputs between models.

Should I pay for several subscriptions at once?

Not necessarily. Map first: find out by testing which tasks are solved best where. Often free models cover part of your tasks — and paying makes sense for the one that owns your main workload. The subscription follows the map, not the other way around.

Channel

Breakdowns and notes — no fluff

New material from Anjela on AI, expertise and marketing. Subscribe to the channel.

Subscribe on WhatsApp