What Can Generative AI Actually Do? A Capability Map for Knowledge Work
A task-level map of where generative AI can transform, retrieve, classify, and propose—and where evidence, judgment, or accountable approval must take over.
Generative AI can draft, transform, classify, extract, compare, simulate, and propose—but none of those verbs guarantees truth or a good decision. Treat capability as a property of a specific task under specific conditions. The strongest uses are reversible and inspectable; the weakest combine missing evidence, ambiguous goals, high consequences, and no accountable reviewer.
Start with the work, not the spectacle
A fluent demonstration invites the wrong question: “How intelligent is this system?” A useful evaluation begins lower down: what operation must be performed, on which material, to what standard, and with what consequence if it fails?
This article is for people deciding where generative AI belongs in research, analysis, writing, learning, or operations. It does not rank vendors or predict when artificial general intelligence will arrive. Model-level rankings age quickly; a task map remains useful because it makes the work and its controls visible.
UNESCO treats AI competence as a combination of human-centred judgment, ethics, techniques, applications, and system design—not mere interface fluency. unesco, nist, helm NIST likewise frames generative-AI risk across the full system and lifecycle. nist Together, they make a crucial distinction: producing an output is not the same as producing a dependable outcome.
The capability and control map
Use four questions to position a task:
- Inspectability: Can a competent person tell whether the output is good?
- Evidence access: Does the system have the material needed to support its claims?
- Reversibility: Can an error be corrected before it creates harm?
- Consequence: What happens if the output is wrong, biased, disclosed, or acted upon?
| Operation | Typical strength | Hidden obligation | Sensible control | |---|---|---|---| | Transform supplied text | Fast restructuring, compression, rewriting | Preserve meaning, qualifications, and provenance | Compare against the source | | Generate options | Breadth and variation | Options may be conventional, impossible, or misframed | Apply explicit constraints and select humanly | | Extract or classify | Repetition at scale | Edge cases and labels may drift | Test a representative set with known answers | | Explain or tutor | Adaptive wording and questioning | Fluency can disguise error or dependency | Ground in trusted material and test unaided performance | | Retrieve and synthesize | Connect evidence across documents | Retrieval can miss, mis-rank, or miscite | Open sources and trace claims | | Recommend or decide | Rapid comparison of stated criteria | Values, missing context, and accountability are not delegated | Keep the decision owner and record reasons | | Act through tools | Executes multi-step work | Permissions magnify mistakes and attacks | Least privilege, checkpoints, logs, and stop rules |
The map does not say “safe” or “unsafe” in the abstract. It shows why summarizing a document you wrote is different from diagnosing a patient, even if both produce a paragraph.
Why the evidence supports the Generative AI Capability and Control Map
NIST identifies risks including confabulation, harmful bias, privacy, information security, and human over-reliance. HELM shows why a single benchmark score is insufficient: model performance must be considered across scenarios and metrics such as accuracy, calibration, robustness, fairness, bias, and efficiency. The evidence supports contextual evaluation, not universal capability rankings.
nist, helmRead capability claims at three levels
Model capability describes what a base model can produce under test conditions. System capability adds retrieval, tools, instructions, memory, interfaces, and safeguards. Operational capability adds real users, incentives, data quality, review time, and consequences.
A model may solve a benchmark problem while an operational workflow fails because the wrong document was retrieved. A system may generate a correct answer while a user cannot distinguish it from an equally fluent false one. Conversely, a modest model can be useful in a tightly bounded transformation task because every change is inspectable.
This yields a better sentence than “AI can write reports”:
In this workflow, the system can produce a first-pass comparison from an approved source set; an analyst checks every decision-relevant claim before the report leaves the team.
That statement identifies the input boundary, operation, reviewer, and release condition.
The boundary-card test
For any proposed use, complete one card:
| Field | Example | |---|---| | Target operation | Compare six vendor policies against five requirements | | Permitted inputs | Current public policies and internal requirement list | | Prohibited inference | Do not infer compliance from missing text | | Success evidence | Correct citations and agreement with expert-coded examples | | Failure threshold | Any invented clause blocks release | | Human owner | Procurement counsel approves the conclusion |
If the team cannot fill in “success evidence” and “failure threshold,” it is not ready to automate. The missing definition is a management problem, not a prompting problem.
Build a capability map for one real workflow
Choose a recurring task and decompose it into operations. Do not map “marketing” or “research”; map “cluster interview excerpts,” “find contrary evidence,” or “rewrite approved copy at three reading levels.”
For each operation:
- Mark the evidence the system may use.
- Rate inspectability, reversibility, and consequence as low, medium, or high.
- Create five routine examples and five edge cases.
- Compare unassisted work with AI-assisted work for accuracy, time, and correction burden.
- Assign an accountable reviewer and a release condition.
- Keep, redesign, or reject the use based on observed results.
Continue with AI Search vs Chatbots vs Agents to choose a system pattern, then apply the four-layer quality model before deployment. If the vocabulary itself is unfamiliar, begin with AI Literacy for Adult Learners.
Category errors that inflate capability
- Judging a consequential workflow from a polished demonstration.
- Treating one benchmark as a universal measure of intelligence.
- Calling retrieval, generation, and verification the same operation.
- Assuming a human reviewer will catch errors without time, evidence, or authority.
- Measuring minutes saved while ignoring correction and coordination costs.
- Generalizing from English, familiar domains, or clean inputs to every user and context.
- Confusing “the system produced an answer” with “the organization can defend the decision.”
Limits of a task-level map
This framework cannot certify a model, product, profession, or regulatory use. Capability changes with model versions, system design, data, language, adversarial pressure, and the evaluator's own expertise. Some harms are difficult to observe in a small test, and some benefits emerge only in sustained use. Before publication or deployment, verify time-sensitive product facts against current provider documentation; for legal, medical, financial, safety-critical, or rights-affecting work, add the relevant qualified authority.
The practical unit of AI capability is not the astonishing answer. It is the bounded operation that remains useful after its evidence, controls, and consequences become visible.
Named sources
Evidence and further reading
Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.