CourionAI
DE
Newsletter
← Blog
opinion 11 Min. Lesezeit

The Most Valuable AI Skill Is Not Prompting. It Is Verification.

A good prompt can improve the first answer. Verification decides whether you should use it. Here is why that distinction matters, what the evidence supports, and a practical way to check AI output without turning every task into an audit.

Risograph illustration of a sunlit wooden desk from above, a printed page with one sentence marked in tangerine highlighter, a pencil line running across a ruler to the same passage circled in an open reference book

Dieser Artikel ist noch nicht auf Deutsch verfügbar. Angezeigt wird die englische Fassung.

Two people ask an AI assistant to research a supplier. The first has taken a prompting course. Their request specifies a role, a format, a tone, a target audience and seven constraints. The answer arrives as a clean comparison table.

The second person writes three ordinary sentences. Then they open the supplier’s price page, check the cancellation terms and follow the link behind the one number that would decide the purchase.

The first person is better at getting an answer. The second is better at knowing whether the answer can be used.

That difference is becoming more important as AI systems get easier to talk to. You no longer need a small incantation to make a capable model summarise a document or rewrite an email. You do need a method for noticing when a polished answer contains an old price, a source that says something else, or a confident sentence supported by nothing at all.

Prompting helps you shape an answer. Verification tells you whether you are allowed to trust it. As models get easier to prompt, verification becomes the more valuable skill.

A better prompt can produce a better wrong answer

Prompting and verification solve different problems.

A good prompt makes your intent legible. It tells the model what you want, what context matters and what a useful result should look like. That is real work. Our own guide to giving AI better context is built around it.

But a prompt is still an instruction to the system producing the answer. It is not evidence that the answer is true.

The US National Institute of Standards and Technology calls invented facts from generative AI “confabulations”, often called hallucinations. Its risk profile says these errors are a natural result of how generative models work. They predict plausible continuations from patterns in data. Plausible and accurate overlap often enough to be useful, but they are not the same category.

NIST also points to the more awkward version of the problem: a model can invent the logic or citation that appears to justify its answer. Asking for sources is good practice. Treating the presence of a blue link as proof is not.

Newer models are better at this than older ones, but “better” is not “finished”. OpenAI’s July 2026 system card for GPT-5.6 still includes a dedicated hallucination evaluation based on conversations that users had flagged for factual errors. OpenAI reports an improvement over GPT-5.5 and warns that the test cases are deliberately difficult, not representative of ordinary traffic. Both parts matter. The model improved, and its maker still treats factual error as something that needs measurement.

This is why a longer prompt cannot settle the issue. You can ask for uncertainty, quotations and citations. You should. The resulting answer may be easier to check. It has not checked itself merely because you requested a section called “Sources”.

The work moved from producing to judging

For most office work, AI does not remove thinking. It moves the thinking to a less visible place.

Microsoft researchers surveyed 319 knowledge workers and collected 936 first hand examples of AI use at work. The study was published at the CHI 2025 conference. People who reported greater confidence in the AI also reported doing less critical thinking. People with more confidence in their own ability reported doing more.

The important finding is not that everybody becomes lazy. The researchers found that the nature of critical thinking changes. Less effort goes into producing a first draft. More goes into verifying information, integrating the response into real work and taking responsibility for the final result.

That last part is easy to miss because generation is visible. You watch a page appear in seconds. Verification looks like opening a second tab, checking a date, running a calculation or noticing that a source discusses the United States while your decision concerns Italy. It feels like friction added after the clever part.

In reality, it is the part that turns generated material into work.

There is a useful warning about trusting that feeling. In a 2025 randomised study by the nonprofit METR, 16 experienced open source developers completed 246 tasks in software projects they knew well. Afterward, they estimated that AI had made them 20 percent faster. The measured result went the other way: with the early 2025 tools used in the study, tasks took 19 percent longer.

That result is narrow. It concerned experienced developers, mature repositories and tools from early 2025. METR said in a February 2026 update that newer agents probably improved productivity, but a follow-up experiment had selection and measurement problems too large for a firm estimate. The responsible lesson is not “AI slows developers down”. It is that experienced people can misjudge the effect of a tool even after using it. Feeling faster is a useful observation, not a measurement.

Verification therefore includes more than fact checking. It also means checking whether the system improved the outcome you care about. Did the reply need fewer corrections? Did the analysis find an issue you would have missed? Did the task actually finish sooner once review time was included? A prompt course cannot answer those questions for you.

The same box is not an independent witness

The easiest verification technique is to ask the AI, “Are you sure?” It is also the weakest.

Research on self-correction has not produced one simple answer. A widely cited 2023 study found that language models struggled to correct reasoning errors without external feedback and sometimes changed correct answers into incorrect ones. That should make us suspicious of the ritual where a model writes an answer, critiques it and then declares the revision verified.

The picture has moved since then. Research published in April 2026 found that two open models, Gemma 3 27B and Qwen 2.5 7B, contained internal confidence signals that helped predict which answers were wrong and which errors the model could correct. The tests covered two specific models and tasks involving question answering and language inference. They are evidence that self-correction can work under some conditions, not evidence that every answer now comes with a reliable internal fact checker.

These findings can sit together. Models may detect some of their mistakes. External feedback still changes the quality of the check.

External does not always mean another person. For code, it can be a test that passes or fails. For arithmetic, it can be a calculator. For a quotation, it is the original page. For a deadline, it is the contract clause rather than a summary of the contract. The common feature is an anchor that does not depend on the model agreeing with its earlier self.

This also explains why expertise still matters. You cannot inspect every sentence with equal effort. Knowing which claim carries the consequence is part of the skill. The date in a birthday invitation deserves less checking than the dosage in a medical summary. Verification is not permanent suspicion. It is allocating attention according to risk.

The strongest case against this

The strongest objection is that the line between prompting and verification is artificial.

A skilled user can ask the model to separate facts from assumptions, link every factual claim, quote the supporting passage, state what it could not confirm, use a calculator and run a second pass against a checklist. That prompt does more than improve style. It builds verification into the workflow.

This objection is right, and it is why “prompting does not matter” would be a bad thesis.

Good instructions reduce the cost of checking. Tools with web search can bring primary sources closer. Document systems can attach page numbers. Code agents can run tests. A well designed process catches more than a person manually reading an unstructured answer from top to bottom.

But those features are scaffolding, not final authority. A page number can point to the wrong passage. A real source can be too old. Three publications can repeat the same unsupported claim. A passing test can cover only the easy path. The user still decides what counts as evidence and whether the check matches the consequence.

There is a second objection. Verification does not scale. If AI produces ten times more work, people cannot inspect every line. Also true. The answer is not to pretend everything was checked. It is to design outputs around what can be checked: claim and source pairs, calculations with inputs, drafts with tracked changes, code with tests, and clear labels for uncertain sections. Where the cost of a mistake is high, generation should stop at the point where qualified review begins.

Prompting helps build that process. Verification defines what the process must prove.

Europe has already put this in the job description

Europe’s most useful contribution to this debate is surprisingly practical.

In June 2026, the OECD and European Commission published an AI literacy framework for schools. It defines AI literacy as knowledge, skills and attitudes that let people understand AI, critically evaluate its outputs and use it ethically and creatively. The definition does not centre on writing better instructions. It centres on judgement.

For workplaces, Article 4 of the EU AI Act requires providers and professional users of AI systems to support the development of AI literacy among the people operating those systems. The rule was amended in July 2026. It no longer demands any specific or “sufficient” level, but the obligation to take measures remains. The Commission says organisations should consider the systems they use, the knowledge of their staff, the context and the risks. National supervision began in August 2026.

That does not mean every small company needs a certified prompt course. The Commission explicitly says there is no required certificate or single training format. It does mean that handing employees a chatbot and a link to its instructions is a thin idea of literacy, especially when the work affects customers, hiring, health, credit or legal rights.

A useful company training would spend less time on clever personas and more on three questions: Which outputs require a check? What counts as an independent source here? Who owns the final decision?

Those questions work in every European language and survive the next model update.

A verification habit that fits the stakes

You do not need to turn every email into an audit. Use three levels.

Low stakes: check the fit. For brainstorming, rewriting or a dinner plan, read the output and ask whether it serves your goal. You are checking usefulness, not truth line by line. If an error would cost five minutes, spend seconds checking.

Factual stakes: check the load-bearing claim. Find the one or two facts that would change your decision. Open the primary source. Confirm the exact passage, date, unit and geographical scope. If the AI provides a quotation, search for the words on the page. If it provides a number, reproduce the calculation. Our guide to making a long PDF answer questions uses this approach with page references.

Consequential stakes: check independently. For medical, legal, financial, safety or employment decisions, use the appropriate professional, official record, test environment or second system of evidence. AI can organise the questions and highlight passages. It should not quietly become the final decision maker because its draft was convenient.

One prompt is still worth keeping:

List the factual claims in your answer that would change my decision. For each one, give the primary source, publication date and exact supporting passage. Mark anything you could not verify.

This does not make the output true. It makes the checking surface smaller. That is what a good prompt should do.

What this means for you. If you are starting with AI, learn one verification habit before collecting prompt templates: open the source behind the most important claim. If you use AI every day, measure the whole task, including correction and review time. If you manage a team, define which outputs need human approval and what evidence that approval requires. The person who can produce ten drafts is useful. The person who knows which draft is safe to send is harder to replace.

What would change my mind

Three observable developments would weaken this thesis.

  1. Independent evaluations showing that automated verification reliably catches factual and reasoning errors across unfamiliar domains, without access to a human expert or an external source of truth.
  2. Organisations teaching prompt technique alone achieving the same real world error rates as organisations that train staff to check sources, calculations and consequences.
  3. Mainstream AI systems consistently identifying which of their own claims require checking, attaching evidence that supports those exact claims, and refusing to fill gaps with plausible text.

Until then, prompting remains useful craft. Verification is the skill that decides whether the craft survives contact with reality.

Sources

Nächster Artikel

Someone Else Is Paying for Your AI, Just Not the Part You Think

The comforting story about your 20 euro AI subscription is that it is sold below cost and the real bill is coming. In 2026 that stopped being true. The subsidy is real, but it sits somewhere else, and it reaches you as rationing rather than as a price rise.

Risograph illustration of a cafe table from above, a hand placing coins on a receipt while a second hand slides a folded stack of banknotes under the tray