Decision Framework

Outcome Leverage Framework

What a person can produce end-to-end with current LLM access, how good it is, and how much value it displaces.

Modeled outcome-level scores across 16 dimensions measuring achievability, output quality, displacement value, and contextual modifiers. Scored 1 to 5 per dimension to produce a Raw Achievability Score, a Displacement Multiplier, and a composite Outcome Leverage Score. These are modeled estimates for current commercial LLM capabilities as of early 2026, not predictions about future models.

For: Workers · Founders · Operators · InvestorsVersion 1.0 · 21 March 2026

01 · The core idea

The Task Compression and Human Advantage Framework asks which work resists automation. This framework asks a different question: what can a person actually produce, end-to-end, right now, using commercially available LLMs?

The unit of analysis is not the task. It is the finished outcome. A signed contract. A deployed website. A proofread manuscript. A working Chrome extension. A marketing email that gets sent. The question is not whether AI can help with parts of the work. The question is whether a person with a $20/month subscription, or free access to a lower-end model, can sit down and produce a complete, usable result.

What outcome leverage means

Outcome leverage is the ratio between what you can produce now and what the same outcome used to cost. A generic services contract used to require an attorney, weeks of coordination, and real money. Now it takes 30 minutes and the output is functionally equivalent for most commercial purposes. That is high outcome leverage.

A novel patent filing can also be generated in 30 minutes, but the output would be shredded by a competent examiner. Same tool, radically different leverage.

What this framework measures

Achievability

Can the LLM produce a finished outcome at all?

Output quality

Is the result actually good enough to use?

Displacement value

What did producing this outcome used to cost in time, money, and dependency chains?

The framework combines all three into a single Outcome Leverage Score that surfaces where current LLM access creates the most real-world value for an individual operator.

What this framework does not measure

It does not measure future capability. It does not predict what models will be able to do in six months. It does not account for fine-tuned or custom-trained systems, specialized agent pipelines, or enterprise toolchains. It scores what a person can do today with commercial off-the-shelf LLM access. The ceiling moves. This framework captures the current floor.

02 · How to read this framework

Score outcomes, not tasks

This framework is applied to finished outcomes, deliverables, and products. Not to individual tasks or subtasks within a workflow. The Task Compression Framework already handles task-level scoring. This framework asks: given that you sit down with the intent to produce a specific thing, how completely can you produce it and how much prior cost does that production displace?

The scoring scale

Each of the 16 dimensions is scored from 1 to 5. A score of 1 means the outcome is difficult to achieve on that dimension (hard to specify, low quality, low displacement). A score of 5 means the outcome is strong on that dimension (easy to specify, high quality, high displacement). The polarity is consistent: higher scores always mean more leverage.

Three scoring layers plus metadata

Layer 1 — Output Achievability

Dimensions 1–4

Whether you can get a finished result in a practical session. Measures how describable the outcome is, whether it completes in one sitting, how many tools you need beyond the LLM, and how much domain knowledge the operator needs.

Layer 2 — Output Quality

Dimensions 5–9

Whether the result is actually good. Structure, accuracy, professional parity, edge case coverage, and taste. A technically completable outcome that produces bad output is not high leverage.

Layer 3 — Displacement Value

Dimensions 10–13

What the outcome used to cost. Money, time, dependency chains, and friction. High displacement is where outcome leverage actually lives.

Decision Metadata

Dimensions 14–16

Modifies interpretation without changing the core score. Frequency, stakes, and audience sophistication determine whether the leverage is practically exploitable.

03 · Layer 1 — Output Achievability

These dimensions measure whether a person can sit down and produce a finished, usable outcome using current LLM access.

01

Specification Clarity

Can you describe what you want in natural language well enough for the model to act? Outcomes that require deep domain knowledge just to articulate the ask score low. Outcomes where any literate person can describe the desired result score high.

1 Requires deep domain expertise to specifyAnyone can describe it 5
02

Single-Session Completability

Can you get a usable output in one sitting? Some outcomes require iteration across multiple days, contexts, or feedback cycles. Others land in one shot or with light refinement within a single session.

1 Many sessions / days / feedback cyclesOne-shot or light iteration 5
03

Tool Chain Simplicity

How many tools beyond the LLM do you need to reach a deployable result? An outcome that requires only the chat window scores high. An outcome that requires a local dev environment, cloud accounts, CI pipelines, and specialized software scores low.

1 Complex toolchain requiredJust the chat window 5
04

Operator Domain Knowledge

How much does the human need to know to prompt effectively and evaluate whether the output is correct? A contract requires knowing what clauses matter. A marketing email requires knowing your audience. Score reflects the knowledge floor, not the ceiling.

1 Deep expertise required to evaluateGeneral literacy sufficient 5

04 · Layer 2 — Output Quality

These dimensions measure whether the result is actually good enough to use, ship, send, or deploy without significant rework.

05

Structural Correctness

Does the output have the right shape, format, sections, and completeness? A contract with the right clauses in the right order. A website with proper routing and responsive layout. An email with appropriate greeting, body, and CTA.

1 Consistently malformed or incompleteStructurally sound out of the box 5
06

Substantive Accuracy

Is the content factually, legally, or technically correct? A contract that includes unenforceable clauses scores low. A codebase that compiles and runs scores high. Medical or legal content that would be dangerous if trusted without review scores 1.

1 Dangerous to trust without expert reviewReliably accurate 5
07

Professional Parity

How close is the output to what a competent professional would produce for the same brief? Not a world-class expert. A solid, competent professional doing commercial work. Score 5 means indistinguishable or better. Score 1 means obviously amateur.

1 Obviously amateurIndistinguishable from professional 5
08

Edge Case Handling

Does the output account for exceptions, gotchas, and real-world variation? A contract that handles termination, IP assignment, and liability. A website that handles empty states and error pages. An app that handles auth failures gracefully.

1 Misses critical edge casesRobust coverage 5
09

Taste and Polish

Does it feel finished, intentional, and well-crafted? Generic boilerplate that reads like a template scores low. Output with voice, design sense, appropriate tone, or craft scores high. This dimension captures the gap between functional and good.

1 Generic / flat / template feelHas voice, design sense, or craft 5

05 · Layer 3 — Displacement Value

These dimensions measure what the outcome used to cost before current LLM access existed. High displacement is where outcome leverage actually lives. An outcome that is easy to produce but was already easy to produce has low leverage.

10

Prior Cost

What did this outcome cost to produce before LLMs? Includes direct fees (attorney, designer, developer), software licenses, and opportunity cost. A services contract that used to cost $2,000–$5,000 in attorney fees scores 5. A marketing email that cost an hour of a copywriter’s time scores 2.

1 Trivial / freeThousands of dollars or more 5
11

Prior Time

How long did the old workflow take from initiation to finished deliverable? Not just active work time. Total elapsed time including coordination, review cycles, scheduling, and waiting. A contract that took 2–4 weeks scores 5. A proofread that took a day scores 2.

1 MinutesWeeks or months 5
12

Prior Dependency Chain

How many other people or organizations were required to produce this outcome? A contract required an attorney, sometimes a paralegal, sometimes opposing counsel review. A full-stack website required a designer, a developer, and a hosting provider.

1 Just youMultiple specialists and vendors 5
13

Prior Path Risk

How much friction, negotiation, failure risk, or rework existed in the old workflow? Miscommunication with vendors, scope creep, revision cycles, dependency on others’ schedules. A contract negotiation that could stall for weeks scores high. An email that just needed writing scores low.

1 Smooth and predictableHigh friction and failure modes 5

06 · Decision metadata

These fields do not describe achievability or quality. They determine whether the leverage is practically exploitable and how much caution the operator should apply.

14

Frequency

How often does someone need to produce this outcome? A marketing email might be weekly. A services contract might be monthly. A full-stack app might be a few times a year. High-frequency outcomes with high leverage compound fastest.

1 Once everWeekly or more 5
15

Stakes

What happens if the output is wrong, incomplete, or bad? A contract with unenforceable terms could cost you a lawsuit. A buggy internal tool wastes time. A bad marketing email gets deleted. Low stakes makes leverage easier to capture. High stakes requires review layers that reduce effective leverage.

1 Catastrophic if wrongLow consequence 5
16

Audience Sophistication

Will the recipient scrutinize the output with domain expertise? A contract reviewed by opposing counsel faces expert scrutiny. A marketing email to a broad list faces general audience evaluation. Expert audiences discount the effective quality of the output.

1 Expert scrutinyGeneral audience 5

Stakes and Audience Sophistication act as discount factors on effective leverage. An outcome with high achievability and high displacement but also high stakes and expert audience scrutiny has less practical leverage than the raw numbers suggest, because the operator must invest additional review, validation, or expert sign-off before the outcome can be used.

07 · Scoring and interpretation

Raw Achievability Score

Add dimensions 1 through 9.

RangeBandMeaning
38–45Reliable one-shotSit down and produce it. Light review, then ship.
28–37Strong with reviewAchievable in one session. Iterate, review, then ship.
19–27Viable draftGets you a real starting point. Needs expert finishing.
9–18Scaffolding onlyProduces structure and fragments. Does not produce a finished outcome.

Displacement Multiplier

avg(D10 … D13) · range 1.0–5.0

Average of dimensions 10 through 13. Captures how much prior cost, time, and dependency the LLM-produced outcome displaces.

Outcome Leverage Score

Raw Achievability × Displacement Multiplier

Surfaces outcomes where high achievability meets high displacement. Maximum possible is 225 (45 × 5.0). The practical ceiling for current models is around 180–200.

Effective Leverage Discount

When Stakes (D15) scores 1 or 2, or Audience Sophistication (D16) scores 1, the operator must add a review layer that reduces the practical speed and cost advantage. This does not change the score. It changes how confidently you can ship the raw output.

Two outputs for every scored outcome

Output 1: Achievability Assessment

Can you produce this, and how good is it?

Output 2: Leverage Assessment

How much prior cost, time, and friction does this displace?

08 · Five bands

Every scored outcome maps into one of five leverage bands. These are the interpretation layer.

AFull Displacement

The LLM produces a finished, usable outcome that displaces a previously expensive or slow workflow. The operator can ship with confidence after light review.

Signs
High achievability, high displacement, low stakes or general audience. The old path involved specialists, weeks, and real money.
Examples
Generic services contracts, marketing emails, Chrome extensions, landing pages, boilerplate legal documents, standard proposals.
What this means
You no longer need the old dependency chain. The skill shifts from production to specification and review.

BHigh Leverage with Review

The LLM produces a strong output that needs targeted review before deployment. The displacement value is high, but stakes or audience sophistication require a validation pass.

Signs
Strong achievability, high displacement, but moderate stakes or some expert scrutiny expected.
Examples
Full-stack web apps, integration-heavy projects, client-facing proposals, internal policy documents, detailed technical documentation.
What this means
You still capture most of the time and cost savings. The review layer is the new bottleneck, not the production.

CViable Draft

The LLM produces a real starting point that accelerates the workflow but does not replace it. Expert finishing is required.

Signs
Moderate achievability, quality gaps in accuracy or edge cases, but the structure and direction are correct.
Examples
Book chapters, patent applications, complex legal filings, enterprise architecture documentation, detailed financial models.
What this means
You get to the 60–70% mark fast. The remaining 30–40% still requires domain expertise and human judgment.

DScaffolding

The LLM produces structure, fragments, and first-pass material. Not a finished outcome. Useful as a starting framework, not as a deliverable.

Signs
Low achievability scores, quality problems across multiple dimensions. The outcome is too complex, too context-dependent, or too novel for current models.
Examples
Novel research papers, complex data platforms, production ML pipelines, custom hardware integration, bespoke enterprise systems.
What this means
The LLM accelerates ideation and initial structure. Production still requires traditional workflows and expertise.

EOut of Reach

The LLM cannot produce a meaningful version of this outcome. The specification complexity, tool chain requirements, or quality demands exceed current capabilities.

Signs
Very low achievability. The outcome requires physical work, real-time environmental interaction, proprietary systems access, or regulated sign-off.
Examples
Surgical procedures, live crisis command, physical construction, courtroom advocacy, hands-on equipment repair.
What this means
These outcomes remain fully human. The LLM may assist with planning, preparation, or documentation around the outcome, but cannot produce it.

09 · Worked examples

Six outcomes scored against all 16 dimensions to demonstrate calibration and illustrate how the framework separates real leverage from superficial capability.

Generic Services Contract (Marketing)

Full Displacement
38/45
Achievability
4.75
Displacement
180.5
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
5553544435545333

High achievability (38/45). The operator needs enough domain knowledge to know what clauses matter (D4=3), but the output is structurally and substantively strong. Displacement multiplier 4.75. Outcome Leverage Score 180.5. Prior path: attorney coordination, weeks, $2,000–$5,000. Current path: 30 minutes.

Marketing Email

High Leverage with Review
40/45
Achievability
2.00
Displacement
80.0
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
5554554432222555

Very high achievability (40/45), though the review needed is minimal. But displacement multiplier is only 2.0, so the Outcome Leverage Score is 80. Copywriters were already fast and cheap. The leverage is real but modest because the prior workflow was not expensive.

Full-Stack Web App (Next.js + Integrations)

High Leverage with Review
29/45
Achievability
4.75
Displacement
137.8
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
4322444335554244

Moderate achievability (29/45). Tool chain scores low (D3=2) because you need a dev environment, package manager, hosting, and DNS. Operator domain knowledge is a real gate (D4=2). But displacement multiplier is 4.75, for an Outcome Leverage Score of 137.8. Prior path: hire a developer, weeks of coordination, $5,000–$15,000. Current path: 150 minutes.

Book Chapter (One-Shot)

Strong with Review
28/45
Achievability
3.25
Displacement
91.0
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
4353432223433132

Achievability 28/45, at the lower edge of Strong with Review. The LLM produces a chapter-length output with correct structure, but professional parity (D7=2) and taste (D9=2) are weak; the output reads like competent filler, not authored prose. Displacement multiplier 3.25. Outcome Leverage Score 91.0. In practice this behaves like a viable draft: you get a strong starting point, not a finished chapter.

Chrome Extension

Strong with Review
34/45
Achievability
4.00
Displacement
136.0
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
4433544434444244

Achievability 34/45. The model handles manifest, content scripts, popup UI, and storage API well. Tool chain (D3=3) requires a text editor and Chrome dev mode. Displacement multiplier 4.0. Outcome Leverage Score 136. Prior path: hire a Chrome extension developer or spend days learning the API yourself.

Synthetic Data Platform

Scaffolding Only
17/45
Achievability
5.00
Displacement
85.0
Leverage
D1D2D3D4D5D6D7D8D9D10D11D12D13D14D15D16
2211322225555121

Achievability 17/45. The specification is hard to articulate (D1=2), requires a complex toolchain (D3=1), and deep operator expertise (D4=1). The LLM can scaffold architecture, generate data models, and write individual components, but cannot produce a working platform. Displacement multiplier 5.0. Outcome Leverage Score 85. The leverage score is deceptive because achievability is too low to capture it.

10 · Relationship to the Task Compression Framework

The Task Compression and Human Advantage Framework and the Outcome Leverage Framework are complementary. They answer different questions about the same underlying shift.

Task CompressionOutcome Leverage
Unit of analysisIndividual task or subtaskFinished outcome or deliverable
Core questionWhich work resists automation?What can you produce end-to-end right now?
Scoring directionHigher = more exposed to disruptionHigher = more leverage for the operator
AudienceStrategic (workforce, investment, policy)Operational (what to do this week)
Time horizonMedium to long termRight now, current models
AssumesTask-level decomposition of rolesOperator at a keyboard with LLM access

Used together: Task Compression tells you which parts of your role are under pressure. Outcome Leverage tells you what you can produce right now to capture value from that pressure. The first framework is strategic. The second is operational.

11 · How to use this correctly

For individuals

List 5 to 10 outcomes you regularly need to produce or that you currently pay others to produce. Score each one. Sort them by Outcome Leverage Score. Start with the highest-leverage outcomes you are not yet producing yourself. The question is not whether you can do everything. The question is where the most value is sitting unattended.

For founders and operators

Score the outcomes your business produces for clients or customers. Where the Outcome Leverage Score is high, your pricing power is under pressure because your customers can increasingly produce the outcome themselves. Where achievability is low but displacement is high, you have a window: offer the outcome at scale before models improve enough for customers to self-serve.

For teams

Score the deliverables your team produces. High-leverage outcomes that are currently bottlenecked on specialists represent the biggest efficiency gains. Low-achievability outcomes that consume significant team time represent the areas where human expertise remains the binding constraint.

What not to do

Do not score hypothetical future capabilities. Score what works today. Do not assume that high achievability means the output requires no review. Always apply the Stakes and Audience Sophistication modifiers before deciding to ship without expert review. Do not confuse scaffolding with finished work. An LLM that produces 70% of a patent application has not produced a patent application.

Do not ask whether AI can produce the outcome. Ask whether you can produce the outcome, right now, with the tools available to you. Then ask what that outcome used to cost.

12 · Compact reference

#DimensionScore 1Score 5
01Specification ClarityDeep domain expertise neededAnyone can describe
02Single-Session CompletabilityMulti-session / multi-dayOne-shot
03Tool Chain SimplicityComplex toolchainChat window only
04Operator Domain KnowledgeDeep expertise to evaluateGeneral literacy
05Structural CorrectnessMalformedSound
06Substantive AccuracyDangerous to trustReliably accurate
07Professional ParityObviously amateurIndistinguishable
08Edge Case HandlingMisses critical casesRobust
09Taste and PolishGeneric / flatCrafted
10Prior CostTrivialThousands+
11Prior TimeMinutesWeeks+
12Prior Dependency ChainJust youMultiple specialists
13Prior Path RiskSmoothHigh friction
14FrequencyOnce everWeekly+
15StakesCatastrophicLow consequence
16Audience SophisticationExpert scrutinyGeneral audience

Dimensions 14–16 are decision metadata that modify interpretation, not the core score.

Find the value sitting unattended

Score your outcomes, capture the leverage that already exists, and build where displacement is highest.

Start a Conversation