79 / KEPNER–TREGOE / SITUATION · PROBLEM · DECISION · POTENTIAL PROBLEM
79 / METHODOLOGY / RATIONAL PROCESS · DIAGNOSIS · DECISION · RISK

KEPNER–
TREGOE.

Kepner–Tregoe (KT) — методология рационального анализа, которая разделяет четыре разных типа управленческой работы: разобраться, что происходит; найти причину отклонения; выбрать вариант; заранее увидеть, что может пойти не так.

Главный принцип: не смешивать diagnosis, decision и risk analysis в один хаотичный разговор. Если причина неизвестна — сначала диагностируем. Если причина известна и нужно выбрать действие — принимаем решение. Если решение уже выбрано — анализируем потенциальные проблемы и готовим preventive/contingent actions.
00. ARCHITECTURAL STATUS

DESIGN-TIME / OPERATIONS METHODOLOGY, NOT A NEW AI AGENT

В нашем большом плане №79 относится к методологическому слою. KT можно применять в architecture reviews, incident analysis, product decisions, release planning, vendor/model selection и postmortem work. Отдельный “KT agent” не нужен: это structured reasoning procedure, которую может исполнять человек, команда или AI-assistant.
TYPEMETHODOLOGYStructured rational analysis.
DEFAULTN/ANot a permanent runtime step.
USE WHENDIAGNOSE / CHOOSE / DE-RISKEspecially under ambiguity and pressure.
SEPARATE COMPONENTN/AProcedure/template, not service.
LIVES INDESIGN-TIME / OPERATIONSProblem solving and decision support.
COMPLEXITYLOW → MEDIUMCan be lightweight or rigorous.
USE AS A FOUR-MODE TOOLKIT
Минимум 80% ценности: correctly classify the task, use Situation Appraisal to split concerns, use IS / IS NOT analysis for unknown causes, separate MUST objectives from weighted WANT objectives for decisions, explicitly score adverse consequences, and perform Potential Problem Analysis on the chosen plan with preventive actions, contingent actions and triggers.
01A. ARCHITECTURE BOUNDARIES & OPERATIONS

ГДЕ KT СТОИТ ОТНОСИТЕЛЬНО СОСЕДНИХ МЕТОДОВ

A. BOUNDARY WITH NEIGHBORS

№78 ТРИЗ сильнее для invention и снятия противоречий; KT сильнее для causal diagnosis, structured choice и implementation risk. №80 TOC ищет системное ограничение; KT помогает диагностировать конкретное отклонение и выбрать действие. №81 IDEF0 моделирует функции системы; KT не моделирует architecture полностью. №82 BPMN описывает поток процесса; KT используется, когда внутри процесса возникла проблема или нужно принять решение.

B. PREREQUISITES / CROSS-REFERENCES

Хорошо сочетается с №13 Problem Formulation, №23 Uncertainty, №25 Evidence-First, №30 What-If, №36 Formal Solvers, №45 Verification, №46 Observability и №47 Evals. Для AI operations KT особенно полезен поверх traces/metrics: факты приходят из observability, KT организует анализ.

C. CONTROL / DATA / RUNTIME / OFFLINE

CONTROL PLANE: templates, scoring rules, priority/risk criteria. DATA PLANE: facts, deviations, IS/IS NOT matrix, objectives, alternatives, risks, preventive actions. RUNTIME: может использоваться on-demand для incident/decision support, но не как default request step. OFFLINE: incident reviews, architecture choices, release readiness, vendor/model selection.

D. FAILURE & OPERATIONS CONTRACT

Success: facts separated from assumptions; cause hypotheses tested against all known facts; decision alternatives evaluated against explicit objectives; implementation risks get owner/action/trigger. Failure: jumping to favorite cause, fake precision in scores, double-counting criteria, ignoring MUST failures, or using risk analysis as excuse never to act.

E. WHAT THIS TOPIC DOES NOT OWN

KT does not prove causality by itself, invent radically new mechanisms, optimize mathematical objective functions, replace experiments, or determine business values. It structures thinking and evidence use; tests, telemetry, domain expertise and evaluation still decide what is true.

01. FIRST DECISION: WHICH KT PROCESS?

НЕ НАЧИНАТЬ “РЕШАТЬ ПРОБЛЕМУ”, ПОКА НЕ ПОНЯТНО, КАКОЙ ЭТО ТИП РАБОТЫ

SITUATION APPRAISAL

Много всего происходит

Несколько concerns, непонятно приоритеты и следующий аналитический шаг.

PROBLEM ANALYSIS

Есть отклонение

Что-то работает не так, как должно, и причина неизвестна.

DECISION ANALYSIS

Нужно выбрать

Есть несколько вариантов и критерии успеха.

POTENTIAL PROBLEM

План уже выбран

Нужно предотвратить или подготовиться к возможным сбоям реализации.

Один incident может пройти через все четыре режима: сначала разобрать ситуацию → найти cause → выбрать fix → de-risk deployment.
02. SITUATION APPRAISAL

РАЗДЕЛИТЬ ХАОС НА УПРАВЛЯЕМЫЕ CONCERNS

LIST CONCERNSЧто требует внимания?
SEPARATEНе смешаны ли несколько проблем?
CLARIFYКакое конкретно deviation/decision/risk?
PRIORITIZESeriousness / urgency / growth.
CHOOSE PROCESSProblem, Decision, Potential Problem?
ASSIGNOwner / next action / deadline.
Situation Appraisal предотвращает типичную ошибку: команда пытается одной root cause объяснить пять разных симптомов, которые на самом деле имеют разные причины и owners.
03. PRIORITY: SERIOUSNESS / URGENCY / GROWTH

НЕ ПУТАТЬ ГРОМКОЕ С ВАЖНЫМ

SERIOUSNESS

Насколько тяжёл эффект?

Revenue, customer impact, safety, legal exposure, SLA, data loss, strategic importance.

URGENCY

Когда надо действовать?

Deadline, active outage, expiring opportunity, irreversible window.

GROWTH

Становится ли хуже?

Error rate rising, backlog compounding, affected tenants expanding, cost accelerating.

CONCERN PRIORITY

Concern: retrieval latency increased

Seriousness:
  medium — p95 misses internal target,
  but no user outage yet

Urgency:
  high — release rollout reaches 100% tomorrow

Growth:
  high — queue depth rising 18% / hour

Action:
  PA immediately
  owner: Search Platform
04. PROBLEM ANALYSIS

“PROBLEM” В KT = ОТКЛОНЕНИЕ ОТ ОЖИДАЕМОЙ НОРМЫ С НЕИЗВЕСТНОЙ ПРИЧИНОЙ

EXPECTED

What should be happening?

ACTUAL

What is happening instead?

DEVIATION

Precise gap whose cause we must explain.

Если причина уже известна — это больше не Problem Analysis. Дальше нужен выбор/реализация corrective action.
05. PROBLEM STATEMENT

НАЗВАТЬ OBJECT + DEVIATION, БЕЗ ПРИЧИНЫ В ФОРМУЛИРОВКЕ

BAD

“Database overload causes slow RAG”

Причина уже встроена в statement и незаметно становится гипотезой по умолчанию.

GOOD

“RAG p95 latency rose from 450ms to 1.8s after 14:20”

Object + deviation + observable facts. Причина пока открыта.

Problem statement должен описывать что не так, а не почему это произошло.
06. THE IS / IS NOT MATRIX

САМАЯ СИЛЬНАЯ ЧАСТЬ KT PROBLEM ANALYSIS

DimensionIS — где deviation наблюдаетсяIS NOT — где могло быть, но не наблюдается
WHATКакой объект? Какая именно deviation?Похожие объекты/отклонения, которых нет.
WHEREГде физически/логически? В какой части объекта?Где рядом/аналогично, но всё нормально?
WHENКогда впервые? Когда повторяется? На каком этапе lifecycle?Когда могло проявиться, но не проявляется?
EXTENTСколько объектов? Насколько велико отклонение? Тренд?Какие масштабы/частоты не затронуты?
IS NOT — не случайный негативный пример. Он должен быть близким сравнением: там проблема могла бы логично возникнуть, но не возникла. Именно различие между IS и IS NOT даёт диагностическую силу.
07. EXAMPLE IS / IS NOT

MODEL LATENCY REGRESSION

PROBLEM:
  Model response p95 rose from 700ms to 2.4s.

WHAT
  IS:
    model profile "standard-v4"
  IS NOT:
    "fast-v2" on same gateway

WHERE
  IS:
    EU deployment
  IS NOT:
    US deployment

WHEN
  IS:
    requests after 14:20 deployment
  IS NOT:
    requests before 14:20

EXTENT
  IS:
    long-context requests > 40k tokens
  IS NOT:
    requests < 10k tokens

DISTINCTIONS:
  standard-v4 + EU + new deployment + long context

CHANGES:
  new serving config at 14:20
  KV-cache policy changed only on EU standard pool

POSSIBLE CAUSE:
  new KV-cache configuration causes additional
  long-context prefill/eviction overhead.

TEST:
  explains WHAT? yes
  WHERE? yes
  WHEN? yes
  EXTENT? yes

NEXT:
  inspect metrics / reproduce / rollback one pool
  to verify causal hypothesis.
08. DISTINCTIONS

ЧЕМ IS ОТЛИЧАЕТСЯ ОТ IS NOT?

CONFIG

Different setting?

Model version, timeout, parser mode, index parameters, rollout flag.

DATA

Different input?

Tenant, document type, language, context size, traffic source.

ENVIRONMENT

Different place?

Region, host class, browser, OS, provider, network path.

TIME / CHANGE

What became different?

Deployment, migration, traffic shift, dependency update, policy change.

Distinction не доказывает cause. Она создаёт сильный список мест, где искать изменения и возможные причины.
09. CHANGES

ЧТО ИЗМЕНИЛОСЬ В ИЛИ ОКОЛО DISTINCTION?

RECENT CHANGE

Near first occurrence

Release, config, dependency, dataset, traffic, provider behavior.

LATENT CHANGE

Old condition, new threshold

Nothing deployed today, but backlog/data volume crossed threshold.

COMBINATION

Two conditions meet

Each existed before, but only together produce deviation.

“Что изменилось?” полезно, но KT не ограничивает диагностику только последним deploy. Причина может быть старой, а проявление начаться после накопления/threshold crossing.
10. POSSIBLE CAUSES

CAUSE MUST EXPLAIN ALL FACTS, NOT JUST ONE SYMPTOM

DISTINCTIONOnly EU pool differs.
+
CHANGEKV config changed.
CAUSE HYPOTHESISNew config harms long-context path.
TEST WHATWhy only standard-v4?
TEST WHEREWhy only EU?
TEST WHENWhy after 14:20?
TEST EXTENTWhy only long contexts?
A candidate cause that needs four ad hoc excuses to fit the IS/IS NOT table is weaker than one that naturally explains the complete pattern.
11. VERIFY THE CAUSE

KT NARROWS HYPOTHESES; EXPERIMENT / TELEMETRY CONFIRMS

OBSERVE

Expected signature

If cause is true, what metric/log/state should differ?

REPRODUCE

Controlled test

Can we recreate deviation by introducing suspected condition?

REMOVE

Rollback / disable

Does deviation disappear when suspected cause is removed?

COUNTERFACTUAL

IS NOT check

Can we explain why comparable unaffected cases remain unaffected?

Do not label “most probable cause” as verified root cause until evidence supports it.
12. MULTIPLE CAUSES

НЕ НАСИЛОВАТЬ ДАННЫЕ РАДИ ОДНОЙ ROOT CAUSE

ONE DEVIATION

One dominant cause

Classic PA works well when a coherent deviation has a common cause.

MIXED SYMPTOMS

Split first

If facts do not fit one cause, Situation Appraisal may reveal multiple problems.

CAUSAL CHAIN

Cause → effect → secondary effect

Distinguish initiating cause from amplifiers and consequences.

13. DECISION ANALYSIS

ВЫБИРАТЬ НЕ “ЛЮБИМЫЙ ВАРИАНТ”, А АЛЬТЕРНАТИВУ ПРОТИВ ЯВНЫХ OBJECTIVES

DECISION STATEMENTWhat exactly are we choosing?
OBJECTIVESMUST and WANT.
WEIGHTSRelative importance of WANTs.
ALTERNATIVESReal candidates.
SCOREEvidence-based rating.
ADVERSE CONSEQUENCESRisk of top alternatives.
DECIDEBest balance + documented rationale.
14. DECISION STATEMENT

ЗАФИКСИРОВАТЬ SCOPE ВЫБОРА

BAD

“Choose the best AI architecture”

Too broad: unclear timeframe, workload, deployment boundary and what counts as an alternative.

GOOD

“Select the primary vector search platform for EU production RAG for the next 18 months”

Clear object, use case, environment and decision horizon.

15. MUST OBJECTIVES

PASS / FAIL. НЕ КОМПЕНСИРУЮТСЯ ВЫСОКИМ SCORE ПО ДРУГИМ КРИТЕРИЯМ

EXAMPLES

Non-negotiable

EU data residency, required API capability, budget ceiling, legal requirement, hard availability requirement.

MEASURABLE

Define threshold

“Fast” is not MUST. “p95 search latency ≤ 150ms at 20M vectors under benchmark X” can be.

SCREEN

Eliminate

Any alternative failing a true MUST is removed before WANT scoring.

ANTI-PATTERN

Everything is MUST

If every preference becomes mandatory, no alternative survives and no trade-off can be evaluated.

16. WANT OBJECTIVES

DESIRABLE, WEIGHTED, TRADEABLE

DECISION:
  primary vector platform

MUST:
  EU deployment available
  tenant filtering supported
  deletion supported
  max annual infra budget ≤ X
  benchmark correctness threshold ≥ Y

WANTS:
  operational simplicity           weight 10
  low p95 latency                  weight 9
  low total cost                   weight 8
  good hybrid filtering            weight 8
  mature backup/recovery           weight 7
  easy observability               weight 6
  low lock-in                      weight 5
Weights express relative importance inside this decision. They are not universal truths about architecture.
17. SCORING

WEIGHT × RATING — USEFUL, BUT DON'T PRETEND IT'S SCIENTIFIC PRECISION

ObjectiveWeightAlt A ratingA weightedAlt B ratingB weighted
Operational simplicity10990660
p95 latency97631090
Total cost8864540
Hybrid filtering8648972
TOTAL265262
A 265 vs 262 score is effectively a tie unless evidence precision supports that distinction. Use scores to make trade-offs visible, then perform sensitivity and risk analysis.
18. RATING QUALITY

SCORE SHOULD POINT TO EVIDENCE

BENCHMARK

Measured

Latency, throughput, cost, recall, availability.

REFERENCE

Known evidence

Vendor docs, architecture constraints, verified case data.

EXPERT JUDGMENT

Subjective but explicit

Mark confidence and rationale instead of disguising opinion as measurement.

UNKNOWN

Do not invent

If critical rating lacks evidence, define experiment/POC rather than guessing a 7/10.

19. SENSITIVITY ANALYSIS

ЕСЛИ НЕБОЛЬШОЕ ИЗМЕНЕНИЕ WEIGHT ПЕРЕВОРАЧИВАЕТ DECISION — РЕШЕНИЕ ХРУПКОЕ

WEIGHTS

Change priorities

Would decision change if cost weight moves from 8 to 9?

RATINGS

Measurement uncertainty

What if latency rating is one point lower than benchmark estimate?

SCENARIOS

Future conditions

What if traffic doubles, model changes, region requirement expands?

Fragile decision → collect better evidence, run POC, or choose an architecture preserving option value.
20. ADVERSE CONSEQUENCES

THE HIGHEST SCORE IS NOT AUTOMATICALLY THE BEST DECISION

WHAT COULD GO WRONG?

Alternative-specific

Migration failure, vendor lock-in, hidden operational burden, data risk, skill shortage.

PROBABILITY

How likely?

Use qualitative/quantitative estimate with evidence and uncertainty.

SERIOUSNESS

How bad?

Assess impact if consequence occurs; consider detectability/reversibility separately where useful.

KT Decision Analysis asks: “Which alternative best meets objectives considering adverse consequences?” — not simply “which row has biggest weighted sum?”.
21. DECISION RECORD

DOCUMENT WHY THE DECISION MADE SENSE AT THE TIME

{
  "decision_id": "DEC-...",
  "statement": "...",
  "date": "...",
  "musts": [...],
  "wants": [
    {"name":"latency","weight":9}
  ],
  "alternatives": [...],
  "evidence_refs": [...],
  "scores": {...},
  "adverse_consequences": [...],
  "selected": "B",
  "rationale": "...",
  "assumptions": [...],
  "review_trigger": "traffic > 2x or vendor pricing changes"
}
WHY

Decision ≠ prediction

A good decision can still have a bad outcome. Record lets you later distinguish poor reasoning from unforeseeable change and avoid hindsight bias.

22. POTENTIAL PROBLEM ANALYSIS

ПОСЛЕ ВЫБОРА: “КАК ИМЕННО ЭТОТ ПЛАН МОЖЕТ ПРОВАЛИТЬСЯ?”

PLANChosen action and steps.
CRITICAL AREASWhere failure matters most?
POTENTIAL PROBLEMSWhat could go wrong?
LIKELY CAUSESWhy might it happen?
PREVENTIVE ACTIONReduce probability.
CONTINGENT ACTIONReduce impact after occurrence.
TRIGGERHow will we know to act?
23. PREVENTIVE vs CONTINGENT

НЕ ПУТАТЬ “НЕ ДАТЬ СЛУЧИТЬСЯ” И “ЧТО ДЕЛАТЬ, ЕСЛИ СЛУЧИЛОСЬ”

PREVENTIVE ACTION

Before failure

Reduce probability: canary, schema compatibility check, capacity headroom, backup verification, permission test, training.

CONTINGENT ACTION

After trigger

Reduce impact: rollback, fallback provider, traffic switch, disable feature, restore backup, human escalation.

A rollback plan is not preventive. It is contingent. Preventive action is what reduces chance you will need the rollback.
24. TRIGGERS

CONTINGENT ACTION БЕСПОЛЕЗЕН, ЕСЛИ НЕПОНЯТНО, КОГДА ЕГО ЗАПУСКАТЬ

METRIC

Threshold

Error rate > 2%, p95 > 1.5s, queue age > 5m.

EVENT

Specific occurrence

Migration checksum mismatch, provider region unavailable.

DEADLINE

No progress by time

If backfill not 90% complete by 18:00, pause cutover.

OWNER

Who acts?

Every trigger needs a person/system responsible for contingent action.

25. EXAMPLE — MODEL ROUTER RELEASE

PPA FOR A REAL AI CHANGE

PLAN:
  route 30% traffic from Model A to Model B.

CRITICAL AREA:
  structured tool-calling.

POTENTIAL PROBLEM:
  Model B emits more invalid tool arguments.

LIKELY CAUSE:
  schema adherence differs on long conversations.

PREVENTIVE:
  run tool-contract eval set;
  validate args deterministically;
  start at 5% canary;
  exclude high-risk tool intents initially.

TRIGGER:
  invalid_tool_args > 0.7%
  OR high-risk schema failure > 0.

CONTINGENT:
  disable Model B route;
  return affected cohort to A;
  preserve traces for failure analysis.

OWNER:
  Model Platform on-call.

SECOND POTENTIAL PROBLEM:
  Model A fallback capacity insufficient.

PREVENTIVE:
  reserve capacity / verify quota.

CONTINGENT:
  degrade optional DEEP work,
  rate-limit noncritical traffic.
26. KT IN INCIDENT RESPONSE

OBSERVABILITY PROVIDES FACTS; KT ORGANIZES THE INVESTIGATION

ALERTS / TRACESRaw symptoms.
SASplit concerns and prioritize.
PAIS / IS NOT, distinctions, changes.
VERIFY CAUSEMetrics/test/rollback.
DAChoose remediation if alternatives exist.
PPARisk of remediation.
DEPLOY / VERIFYObserve expected recovery.
Under incident pressure, KT's value is cognitive hygiene: stop mixing symptoms, causes, solutions and fears in the same sentence.
27. EXAMPLE — RAG QUALITY DROP

PA BEFORE “LET'S CHANGE THE EMBEDDING MODEL”

DEVIATION:
  grounded-answer pass rate fell 91% → 78%.

IS / IS NOT

WHAT IS:
  product manuals corpus
WHAT IS NOT:
  policy documents corpus

WHERE IS:
  documents ingested after parser v3
WHERE IS NOT:
  old indexed documents

WHEN IS:
  after Aug 28 reindex
WHEN IS NOT:
  before reindex

EXTENT IS:
  tables-heavy PDFs
EXTENT IS NOT:
  text-heavy PDFs

DISTINCTIONS:
  parser v3 + new reindex + tables

CHANGES:
  table extraction logic changed
  chunk boundaries changed

CANDIDATE CAUSES:
  A) embedding model degraded
  B) parser table extraction loses row associations
  C) query router changed

TEST AGAINST FACTS:
  A fails WHEN/WHERE distinction:
    same embedding used old documents successfully.
  C fails WHAT:
    same routing across corpora.
  B explains all known IS/IS NOT facts.

VERIFY:
  compare Document IR;
  reconstruct affected table;
  reparse with v2;
  run eval.

LESSON:
  don't optimize vector search before diagnosing
  whether source representation is broken.
28. KT + TRIZ

DIAGNOSE → INVENT → CHOOSE → DE-RISK

KT PROBLEM ANALYSIS

Finds what actually causes the deviation.

TRIZ

If correction faces a hard contradiction, generate inventive mechanisms.

KT DECISION + PPA

Select viable concept against objectives and de-risk implementation.

Это сильная совместная схема: KT tells you what problem you really have; TRIZ helps create a non-obvious way out; KT helps choose and implement it rationally.
29. KT + TOC

BOTTLENECK LOCATION AND LOCAL DIAGNOSIS ARE DIFFERENT QUESTIONS

TOC

Which constraint governs throughput?

Find leverage point in entire flow.

KT

Why did this constraint/deviation change?

Diagnose a specific unexpected behavior, choose corrective action, analyze implementation risks.

TOC can tell you not to waste time optimizing a non-constraint; KT can then investigate why the actual constraint is underperforming.
30. KT + AI

AI GOOD AT ORGANIZING EVIDENCE — DANGEROUS WHEN IT FILLS MISSING FACTS

GOOD

Structure messy incident data

Extract concerns, draft IS/IS NOT matrix, summarize changes and evidence.

GOOD

Generate hypotheses

Propose possible causes linked to distinctions and changes.

CAUTION

Unknown facts

LLM must mark missing cells as unknown rather than “completing” a plausible story.

NO

Cause by confidence

Model saying “most likely” is not verification. Test against telemetry/experiments.

31. AI-ASSISTED PROBLEM ANALYSIS CONTRACT

FORCE FACT / INFERENCE / UNKNOWN SEPARATION

{
  "problem_statement": "...",
  "facts": [
    {
      "dimension":"WHEN",
      "side":"IS",
      "statement":"first seen at 14:20",
      "evidence_ref":"trace://..."
    }
  ],
  "unknowns": [
    "whether US pool received same config"
  ],
  "distinctions": [...],
  "changes": [...],
  "possible_causes": [
    {
      "cause":"...",
      "explains":["WHAT","WHERE","WHEN"],
      "fails_to_explain":["EXTENT"],
      "confidence":"LOW",
      "verification_test":"..."
    }
  ]
}
KEY RULE

Unknown stays unknown

AI assistant may suggest what to measure next, but must not convert absence of evidence into a fabricated IS/IS NOT fact.

32. AI-ASSISTED DECISION CONTRACT

MAKE ASSUMPTIONS AND EVIDENCE AUDITABLE

{
  "decision_statement": "...",
  "musts": [
    {
      "criterion":"EU residency",
      "threshold":"required"
    }
  ],
  "wants": [
    {
      "criterion":"operational simplicity",
      "weight":10
    }
  ],
  "alternatives": [
    {
      "name":"A",
      "ratings":[
        {
          "criterion":"operational simplicity",
          "rating":8,
          "evidence_ref":"..."
        }
      ]
    }
  ],
  "sensitivity_checks": [...],
  "adverse_consequences": [...],
  "selected":"...",
  "assumptions":[...],
  "review_triggers":[...]
}
33. BIAS CONTROL

KT IS USEFUL BECAUSE IT EXTERNALIZES THE REASONING STRUCTURE

ANCHORING

First cause

IS/IS NOT forces candidate cause to fit multiple facts.

CONFIRMATION

Only supporting evidence

Ask explicitly which IS/IS NOT facts candidate cannot explain.

SOLUTION FIXATION

Favorite tool/vendor

Decision objectives defined before rating alternatives.

OPTIMISM

Best-case plan

PPA makes potential failures and triggers explicit before rollout.

KT does not eliminate bias; it creates places where assumptions can be challenged visibly.
34. ANTI-PATTERNS

КАК KT ПРЕВРАЩАЕТСЯ В ТАБЛИЦУ РАДИ ТАБЛИЦЫ

CAUSE IN PROBLEM STATEMENT
Hypothesis becomes assumed fact.
OBJECT + DEVIATION ONLY
RANDOM IS NOT
Comparison has no diagnostic value.
CLOSE COMPARABLE CASE
LAST CHANGE = CAUSE
Post hoc reasoning; threshold/interaction causes ignored.
TEST ALL FACTS
ONE ROOT CAUSE FOR EVERYTHING
Mixed concerns are forced into one story.
SPLIT VIA SA
EVERY CRITERION = MUST
No alternative can pass.
TRUE NON-NEGOTIABLES ONLY
WEIGHTED SCORE = TRUTH
Subjective ratings masquerade as precision.
EVIDENCE + SENSITIVITY
RISK LIST WITHOUT ACTION
PPA becomes anxiety catalog.
CAUSE / PREVENT / CONTINGENCY / TRIGGER
AI FILLS GAPS
Plausible synthetic facts corrupt diagnosis.
UNKNOWN IS A VALID VALUE
35. QUALITY METRICS

HOW TO KNOW THE METHOD IS HELPING

SEP

Concern Separation

How many mixed issues are decomposed into independently actionable concerns?

FIT

Cause Fit

Does verified cause explain all critical IS/IS NOT facts?

TTV

Time to Verified Cause

Incident start → experimentally/observationally confirmed cause.

EVD

Evidence Coverage

% decision ratings backed by measurement/reference rather than unsupported opinion.

ROB

Decision Robustness

Does selected option survive plausible weight/rating changes?

PRE

Prevented Failures

Potential problems caught before implementation through preventive actions.

MTTR

Recovery Readiness

Trigger → contingent action execution time.

RWR

Rework Reduction

Less repeated diagnosis/decision reversal caused by undocumented assumptions.

36. 20-MINUTE PROBLEM ANALYSIS

LIGHTWEIGHT KT FOR INCIDENTS

TimeAction
0–2 minWrite object + deviation without assumed cause.
2–7 minFill WHAT / WHERE / WHEN / EXTENT for IS and closest meaningful IS NOT.
7–10 minList distinctions between IS and IS NOT.
10–12 minList changes around distinctions and first occurrence.
12–16 minGenerate possible causes from distinctions × changes.
16–18 minTest each cause against all known facts. Reject weak fits.
18–20 minDefine fastest safe verification experiment for top cause.
37. 30-MINUTE DECISION ANALYSIS

FOR ARCHITECTURE / VENDOR / MODEL CHOICES

TimeAction
0–3 minWrite precise decision statement and horizon.
3–8 minDefine true MUST objectives with measurable thresholds.
8–13 minDefine WANT objectives and relative weights.
13–18 minScreen alternatives against MUSTs.
18–23 minRate surviving alternatives against WANTs with evidence refs.
23–26 minRun sensitivity: weights/ratings/scenarios.
26–30 minList adverse consequences of top candidates and make decision.
38. PRACTICAL WORKSHEET

COPY FOR REAL WORK

KEPNER–TREGOE WORKSHEET
===========================================================

A. SITUATION APPRAISAL

Concerns:
1.
2.
3.

For each:
Seriousness:
Urgency:
Growth:
Owner:
Next process:
[ ] Problem Analysis
[ ] Decision Analysis
[ ] Potential Problem Analysis


B. PROBLEM ANALYSIS

Problem statement:
OBJECT:
DEVIATION:

WHAT
IS:
IS NOT:

WHERE
IS:
IS NOT:

WHEN
IS:
IS NOT:

EXTENT
IS:
IS NOT:

DISTINCTIONS:
1.
2.
3.

CHANGES:
1.
2.
3.

POSSIBLE CAUSES:
A.
B.
C.

TEST EACH CAUSE AGAINST:
WHAT:
WHERE:
WHEN:
EXTENT:

MOST PROBABLE:
→

VERIFICATION TEST:
→


C. DECISION ANALYSIS

Decision statement:
→

MUST OBJECTIVES:
1.
2.
3.

WANT OBJECTIVES:
criterion / weight
1.
2.
3.

ALTERNATIVES:
A.
B.
C.

SCORES:
→

SENSITIVITY:
→

ADVERSE CONSEQUENCES:
A →
B →
C →

SELECTED:
→

RATIONALE:
→

REVIEW TRIGGER:
→


D. POTENTIAL PROBLEM ANALYSIS

Plan / step:
→

Potential problem:
→

Likely cause:
→

Preventive action:
→

Trigger / early warning:
→

Contingent action:
→

Owner:
→
39. AI PROMPT — PROBLEM ANALYSIS

USE AI WITHOUT LETTING IT INVENT THE INCIDENT

You are assisting with Kepner–Tregoe Problem Analysis.

Rules:
1. Do not propose a root cause immediately.
2. Restate the problem as OBJECT + DEVIATION.
3. Separate facts, inferences and unknowns.
4. Build IS / IS NOT across:
   WHAT, WHERE, WHEN, EXTENT.
5. IS NOT must be a meaningful comparable case.
6. Never invent a missing fact.
7. Identify distinctions between IS and IS NOT.
8. Identify known changes associated with distinctions.
9. Generate candidate causes from distinctions and changes.
10. For every candidate, test whether it explains
    every important IS / IS NOT fact.
11. Explicitly list facts the cause fails to explain.
12. Rank hypotheses only provisionally.
13. Define the smallest safe observation/experiment
    that could verify or falsify the top hypothesis.
14. Do not call a hypothesis a root cause
    until verification evidence exists.
40. AI PROMPT — DECISION ANALYSIS

MAKE THE DECISION STRUCTURE EXPLICIT BEFORE SCORING

You are assisting with Kepner–Tregoe Decision Analysis.

1. Write a precise decision statement.
2. Ask for the decision horizon and scope.
3. Separate objectives into:
   MUST — non-negotiable pass/fail thresholds;
   WANT — desirable weighted objectives.
4. Challenge every MUST:
   is it genuinely mandatory?
5. For every WANT:
   define measurable meaning where possible.
6. List alternatives without scoring yet.
7. Screen alternatives against MUSTs.
8. Score surviving alternatives against WANTs.
9. Cite evidence for each important rating.
10. Mark unsupported ratings as assumptions/unknowns.
11. Run sensitivity on weights and uncertain ratings.
12. Identify adverse consequences for top alternatives.
13. Recommend only after objectives, evidence and risks are visible.
14. Record assumptions and review triggers.
41. PRACTICAL DECISION

WHERE KT EARNS ITS PLACE

QuestionAnswer
Стоит ли использовать?Да. Особенно для incidents, root-cause work, architecture/vendor choices и rollout risk.
Separate Component?N/A. Это методология/процедура, а не отдельный runtime service.
Минимум 80% ценности?Choose correct KT process; IS/IS NOT for diagnosis; MUST/WANT + evidence + sensitivity for decisions; prevention/contingency/trigger for rollout risk.
Когда overkill?Очевидный баг с уже доказанной причиной, тривиальный выбор, либо ситуация, где нужен творческий invention rather than diagnosis.
Trigger?“We have facts but are arguing about the cause”; “we have options but criteria are implicit”; “we chose a plan but rollout risk is unclear”.
Как измерить uplift?Time-to-verified-cause, fewer repeated incidents, decision reversals, evidence coverage, rollout failure rate and recovery readiness.
Можно ли использовать AI?Да. AI is strong at structuring facts/hypotheses/criteria; it must never fabricate missing IS/IS NOT facts or substitute confidence for causal verification.
42. DESIGN RULES

THE KT RULEBOOK

RULE 01

Classify the task first

Situation, cause, choice and future risk require different analysis.

RULE 02

Problem = deviation, not presumed cause

Keep hypotheses out of the problem statement.

RULE 03

IS NOT must be comparable

Diagnostic power comes from meaningful contrast.

RULE 04

Cause must explain all key facts

Reject hypotheses that fit only the loudest symptom.

RULE 05

Verify causality externally

KT narrows search; experiment/telemetry confirms.

RULE 06

MUST before WANT

Non-negotiables screen alternatives before weighted trade-offs.

RULE 07

Scores need evidence

Do not use numerical tables to hide subjective guesses.

RULE 08

Analyze adverse consequences

Highest utility score may still have unacceptable risk.

RULE 09

Every critical risk needs a trigger

Prevent before; contingency after; owner always.

43. FINAL MAP

KT TURNS “WE NEED TO THINK” INTO FOUR DIFFERENT ANALYTICAL JOBS

START
  ↓
WHAT KIND OF SITUATION IS THIS?

────────────────────────────────────────────

MANY CONCERNS / CHAOS
  ↓
SITUATION APPRAISAL

list concerns
  ↓
separate mixed issues
  ↓
clarify
  ↓
priority:
  seriousness
  urgency
  growth
  ↓
owner + next process

────────────────────────────────────────────

UNEXPECTED DEVIATION
CAUSE UNKNOWN
  ↓
PROBLEM ANALYSIS

EXPECTED
vs
ACTUAL
  ↓
OBJECT + DEVIATION
  ↓
IS / IS NOT

WHAT
WHERE
WHEN
EXTENT
  ↓
DISTINCTIONS
  ↓
CHANGES
  ↓
POSSIBLE CAUSES
  ↓
TEST EACH CAUSE
AGAINST ALL FACTS
  ↓
MOST PROBABLE HYPOTHESIS
  ↓
VERIFY WITH:
  telemetry
  experiment
  reproduction
  rollback
  direct observation
  ↓
CAUSE VERIFIED

────────────────────────────────────────────

CAUSE KNOWN / NEED TO CHOOSE ACTION
  ↓
DECISION ANALYSIS

decision statement
  ↓
MUST objectives
  pass / fail
  ↓
WANT objectives
  weighted
  ↓
alternatives
  ↓
MUST screening
  ↓
evidence-based scoring
  ↓
sensitivity analysis
  ↓
adverse consequences
  ↓
select alternative
  ↓
record assumptions
and review trigger

────────────────────────────────────────────

PLAN CHOSEN
  ↓
POTENTIAL PROBLEM ANALYSIS

critical step
  ↓
what could go wrong?
  ↓
likely cause
  ↓
PREVENTIVE ACTION
  reduce probability
  ↓
EARLY WARNING / TRIGGER
  ↓
CONTINGENT ACTION
  reduce impact
  ↓
owner

────────────────────────────────────────────

WITH TRIZ:

KT PA
  finds the real cause/conflict
        ↓
TRIZ
  invents stronger mechanism
        ↓
KT DA
  chooses between concepts
        ↓
KT PPA
  de-risks rollout
        ↓
EVAL / PROTOTYPE / PRODUCTION
  verifies real outcome

────────────────────────────────────────────

IN AI SYSTEMS:

OBSERVABILITY
  provides facts

KT
  organizes diagnosis

TRIZ
  helps invent non-obvious redesign

EVALS
  measure whether change helped

PRODUCTION ARCHITECTURE
  implements safely

────────────────────────────────────────────

THE MOST IMPORTANT KT QUESTION
IS OFTEN NOT:

“WHAT IS THE ROOT CAUSE?”

IT IS:

“WHAT EXACTLY IS DIFFERENT
BETWEEN WHERE THE PROBLEM
IS
AND WHERE IT LOGICALLY
COULD HAVE BEEN
BUT IS NOT?”

THAT CONTRAST
TURNS A WALL OF POSSIBLE CAUSES
INTO A MUCH SMALLER
TESTABLE SET.

ECC RETROFIT / PRACTICAL HARNESS INTEGRATION

A. Related ECC ideas. Context-as-cache, scoped memory, lifecycle hooks, selective capabilities, feature flags, deterministic enforcement, provider-neutral adapters and eval-gated learning are applied only where relevant to №79 Kepner–Tregoe.

B–E. Existing boundary and placement. The existing conceptual boundary, class METHODOLOGY, default N/A and owner Design-time / Problem Solving remain authoritative. Runtime/control/data/offline placement is unchanged; durable state stays outside model context.

F–H. Hooks and contracts. Use bounded PRE_MODEL/POST_MODEL, PRE_TOOL/POST_TOOL, CHECKPOINT and TASK_COMPLETED events as applicable. Illustrative fields and canonical contracts are defined in NEW_CONTRACTS_SPEC.md; no universal schema is implied.

I–J. Security and evaluation. Host-side schema, permission, secret, budget, idempotency and audit checks take precedence over LLM output. Optional mechanisms require a feature flag and WITH/WITHOUT ablation; measure quality, acceptance, correction, latency, cost, escalations and severe errors.

K–L. Task profiles and cross-references. A TaskProfile selects the relevant skill, tool/context slice, memory scope and enforcement profile independently from FAST/STANDARD/DEEP. See cross-reference map, hook spec and ablation plan. Provider adapters remain outside the core.