Coding

Probabilistic Thinking

Try it

Replace vague hunches with calibrated probability estimates you can track and improve over time.

What it does

This skill replaces binary will/won't thinking with a number anchored in historical base rates, updated by evidence using Bayes' theorem, and scored after the outcome resolves. It produces a structured probability estimate with a confidence interval, a named next piece of evidence that would move the estimate, and a calibration log entry for later scoring. Works for any uncertain forecast: deal closes, product launches, hiring outcomes, AI timelines.

When to use it

  • Making a forecast about an uncertain outcome
  • Evaluating a confidence claim like 'I'm 90% sure'
  • Converting a vague question into a specific probability question
  • Replacing a compelling story with a base-rate check

The skill document

Probabilistic Thinking

Overview

Most reasoning is binary: will it happen, or won't it? That framing discards the most useful information — the degree of confidence — and produces predictions that cannot be checked, updated, or scored. Probabilistic thinking replaces binary with calibrated probability estimates: numbers anchored in base rates, updated with evidence, and scored after the fact. Rooted in Bayes (1763), Knight's risk-vs-uncertainty distinction (1921), and Tetlock's empirical work showing calibration is a trainable skill.

Composable neighbors: first-principles · occams-razor · second-order-thinking · inversion · regret-minimization · expected-value-and-kelly. This skill is the upstream input the others depend on — the probability estimate here feeds EV-Kelly, calibrates inversion's failure-path weights, and gives second-order's hops their confidence decay.

When to Use

Use when reasoning about an uncertain outcome (forecast, diagnosis, pipeline conversion, hire, deal close, geopolitical event); when binary "will/won't" predictions are being made; when a vivid story is replacing a base rate; when "I'm 90% sure" appears with no calibration evidence; when forecasting AI timelines / AGI arrival / agentic reliability, or judging whether AI capex, AI valuations, or AI adoption rates justify a point-estimate bet amid genuine uncertainty.

When NOT to use: deterministic problems (math, well-defined engineering); pure Knightian uncertainty with no usable base rate (give a range + humility statement instead); decision is robust across all likely probabilities; question is identity/ethics/meaning (→ regret-minimization).

Coaching Novices (Adaptive Front Door)

  • Engine mode: user has a concrete forecasting question → run The Process directly.
  • Coach mode: user is unfamiliar or has no concrete case → guide step by step.

In Coach mode, respond one step at a time. Each [WAIT] is a hard stop — output only that step's question, then stop.

  1. What-it-is. Probabilistic thinking replaces "will it happen or not" with a number (0–1) anchored in base rates, updated with evidence, and scored after the fact.
  2. Check fit. Match against When to Use / When NOT to use. Redirect if deterministic or pure Knightian.
  3. Elicit their real question. "Odds of success" is vague; "probability customer X signs by Q3 given yesterday's call" is a question.

[WAIT — do not advance until user responds]

  1. Walk The Process one step per turn. Base rate first, then evidence, then update with them.

[WAIT — do not advance until user responds]

  1. Close. State the probability number and the one piece of evidence that would move it most.

[WAIT — do not advance until user responds]

The Process

Run the Probability Estimate. Base rate first, then evidence, then update, then calibration check.

  1. Precise question + deadline. "Will the deal close?" → "Will customer X sign ≥$50K by 2026-09-30?"
  2. Anchor in a base rate. Historical fraction of similar situations. No base rate = Knightian territory → report range, not point.
  3. Evidence for and against. Each signal moves estimate ↑ or ↓. Be uncharitable about both sides.
  4. Bayesian update (plain language). For each signal: P(evidence | outcome happens) vs P(evidence | doesn't happen). The ratio drives the shift.
  5. Number + confidence interval. Not "70-ish" — "68%, 80% CI 55–80%."
  6. Most-informative next evidence. If nothing would move your estimate, you have a belief, not an estimate.
  7. Calibration log. Record estimate, date, resolution criteria. Score after: did 70%-calls land 70% of the time?

Output: the Probability Estimate

Question (precise): 
Base rate:  →  (source, n=)
Evidence:  ↑/↓ strong/moderate/weak  [repeat per signal]
Bayesian shift: net <↑/↓ to X%> — rationale in one paragraph
Estimate:   |  80% CI: <%–%>  |  Knightian caveat if needed
Next evidence:  → <% if X> / <% if Y>
Calibration log: date | question | resolution criteria | Brier score after

→ Method in Action: Tetlock, IARPA, and the Good Judgment Project (2011–2015) → 2026 lens: Forecasting AI Timelines and Agentic Reliability (2023–2026)

Calibration Packs

DomainBase rate sourceClassic failure
MedicalDisease prevalence in populationBase-rate neglect → false positives
SalesConversion-by-stage historyAnchoring on preferred deal
LegalCrime/suspect-pool frequenciesProsecutor's fallacy
Product/startupCohort retention, vintage distributionsSurvivorship bias

Applying It Well

  • Base rate first. A story without a base rate is fiction with a number attached.
  • Incremental updates. Superforecasters update more often but less drastically than amateurs.
  • Risk ≠ uncertainty (Knight 1921). For Knightian situations give a range + humility statement, not false precision.
  • Name what would change your mind. If nothing would move your estimate, you have a position, not an estimate.
  • Score yourself. Write forecasts down and check them — calibration is trainable only this way.

→ Primary sources: references/sources.md

Common Rationalizations

[D] = designed upfront | [O] = observed in real use. [O] entries are more valuable.

Fake moveReality
[D] Base-rate neglectReaching for a vivid story while ignoring the prior frequency of the outcome class. Always anchor in a base rate first; updates come from there.
[D] Confusing P(evidence | outcome) with P(outcome | evidence) (prosecutor's fallacy)A test triggering on 99% of cases of a rare disease will, in a low-prevalence population, produce mostly false positives. Bayes' theorem connects them; they are not interchangeable.
[D] Treating "very likely" as a binary"I'm 90% sure" with no calibration history and no number for "what would make it 50%" is a vibe, not an estimate.
[D] Confusing risk with uncertainty (Knight 1921)An actuarial-table problem and a geopolitical-forecasting problem differ in kind. Inventing a precise number for genuine Knightian uncertainty manufactures false confidence.
[D] One-shot probability fallacy"The probability of this event is X" implicitly invokes a reference class. Name it; otherwise the probability is undefined.
[D] Survivorship bias in base-rate constructionReasoning from winners without including losers gives an inflated base rate. The reference class must include the failures.
[D] Anchoring on the first number that appearsEven random numbers shift estimates (Tversky & Kahneman 1974). Notice when you are anchoring on the latest news rather than the base rate.
[D] Under-updating on strong evidenceStubbornly holding the prior when new information is high-quality. Bayes says update; ignoring evidence is anti-Bayesian.
[D] Over-updating on weak evidenceLetting noisy or single-source data dominate. Superforecasters' edge is smaller updates more often, not bigger ones.
[D] Pseudo-precision"73.2% probability" when inputs justify nothing tighter than "60–80%." Match precision to evidence strength.
→ Add [O] entries here after each real use — paste the actual failure patternWhat went wrong and why

Red Flags

  • Point estimate with no base rate · "high probability" with no number · no evidence named that would change the estimate · single story doing all the work · reference class silently chosen to favor a conclusion · Knightian situation with no humility statement · same person makes many forecasts but has never scored them

Verification

  • The question is stated with a specific outcome and a deadline
  • A base rate is named with an explicit reference class and a source
  • Evidence is listed in both directions, with direction and strength tags
  • The Bayesian shift from base rate is explained in one paragraph
  • The point estimate is a number, with an 80% confidence range
  • The single most-informative next piece of evidence is named
  • The estimate is recorded with date and unambiguous resolution criteria for later scoring

Part of deciqAI Knowledge Skills — 227 open-source thinking skills that make rigor executable for AI agents. The same skills power every deciqAI agent, which runs them autonomously to operate your company. See it run → https://www.deciqai.com/c/probabilistic-thinking · ⭐ Star the repo → https://github.com/deciqAI/knowledge-skills · Contributions welcome.

Agents: latest version & machine-readable metadata → https://www.deciqai.com/s/probabilistic-thinking.json

Questions people ask

What's the difference between this and just guessing?
A guess has no base rate, no evidence list, and no way to improve. This process forces you to anchor in historical frequencies, list evidence in both directions, and record the estimate so you can score yourself later. Calibration is trainable only when you write forecasts down and check them.
When should I use numbers instead of just thinking 'probably'?
Use numbers when a decision hinges on the probability — when you'd act differently at 30% vs 70%. If the outcome doesn't affect any decisions, a probability estimate may be unnecessary overhead.
What if there's no historical data to anchor on?
If no usable base rate exists, you're in Knightian uncertainty territory. Report a wide range and a humility statement rather than manufacturing false precision. The skill flags this situation explicitly.

Related skills

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Make irreversible life decisions by projecting to 80 and naming which regret you'd rather live with.

by deciqai1 installs2 stars

Diagnose feedback loops to predict how complex systems behave and find where intervention actually works.

by deciqai1 installs3 stars

Tests whether you genuinely understand something or just recognize it — exposes the gaps in your mental model.

by deciqai1 installs2 stars

Detect when presentation language is steering your decision instead of the facts themselves.

by deciqai1 installs2 stars

Audit strategies and portfolios for hidden assumptions that break under extreme events.

by deciqai1 installs3 stars

More from deciqai

Browse all skills

Diagnose your organization's strategic phase and spot misaligned initiatives before they drain momentum.

by deciqai7 installs2 stars

When multiple explanations all fit the evidence, pick the one that assumes the least.

by deciqai5 installs2 stars

Identify catastrophic failure paths before you commit — and design your plan around eliminating them.

by deciqai4 installs2 stars

Trace decisions past the obvious effect to catch the consequences that reverse it.

by deciqai4 installs2 stars

A structured interview and analysis process that uncovers the actual job customers hire your product to do — and who they're really competing against.

by deciqai3 installs3 stars

Diagnose why learners are struggling and redesign instruction to stay within working memory limits.

by deciqai2 installs3 stars