The four teaching datasets that accompany The Defensible Decision. Each is sized to support real practice across the 14 chapters — not a toy small enough to fit on a single slide, not so large that loading it becomes a chapter exercise on its own. Each includes the kinds of imperfections real data has (null values, naming inconsistencies, occasional outliers, definitional drift) so the chapter exercises teach the discipline the book is actually about.
All four datasets are wholly synthetic. Any resemblance to a real company's data is coincidental. Each dataset's README documents what it represents, what it does not represent, and the deliberate imperfections built into it. Appendix E provides a completed exemplar and an intentionally underspecified starter for each dataset below, plus the AIRS research exemplar.
The four datasets
CloudRevenue
Cloud subscription revenue across 8 product families × 247 tenants × 36 months. Supports Chapters 3, 4, 5, 9, 10, and 13. Includes region-naming inconsistency, in-flight billing close, and 12 mid-window churners as the canonical practice imperfections.
M365Marketing
Multi-channel campaign performance across 24 campaigns × 5 channels × 78 weeks. Supports Chapters 3, 5, 10, 11, and 12. Includes mixed cross-channel attribution windows (the canonical exercise for Ch 11), paused weeks, and 3 platform-outage diagnostics for anomaly-detection work.
SupportInsights
Customer support tickets at full ticket grain. 47 agents across 3 pods, 10 categories, 60 knowledge-base articles, ~25,400 tickets over 12 months. Supports Chapters 4, 10, 11, and 13. Embeds an APAC-pod Q2 staffing-shortage signature for anomaly detection, a CAT-FEAT / CAT-PROD definition-change for Ch 11 reconciliation work, and ~6% null CSAT for realistic completeness analysis.
EnterpriseGovernance
Data-governance snapshot across 12 business units × 8 governance capabilities. The dataset's distinctive feature: the self-assessment input disagrees systematically with the operational-signal input, which is exactly why the dashboard the starter BRD describes needs to exist. Supports Chapters 11, 13, and 14.
Chapter worked-example datasets
Compact, purpose-built datasets that back a single chapter’s worked example. Unlike the four cross-chapter teaching datasets above, these do not carry designed-in imperfections and are sized to exactly what the chapter’s figures visualise.
Priya’s Q3 sales
The synthetic apparel-retailer scenario behind Priya’s Q3 dashboard in Chapter 1. Aggregate rollups only: three monthly totals, four regional revenue-and-plan pairs, six product lines with revenue and gross-margin decomposition, and one Q2 baseline — all reconciling to a $42.7M Q3 total. Backs the flawed and corrected dashboards on the Chapter 1 companion page.
Aisha’s AIRS cohorts
The synthetic six-cohort scenario behind Aisha’s AIRS readiness-gain chart in Chapter 6. Start-of-course and end-of-course mean AIRS scores across six semester cohorts (Fall 2023 through Spring 2026), with the Cohort 6 gain locked at roughly 2.2× Cohort 1 so the most-recent-vs-first contrast the Big Idea rests on is visible in the chart. Backs the flawed and rebuilt figures on the Chapter 6 companion page.
Rohan’s services dashboard
The synthetic services-company Q3 scenario behind Rohan’s executive dashboard in Chapter 12. Three monthly totals with actual and plan (Q3 opened behind pace, recovered by September), four service lines, four regions, one Q2 baseline — all reconciling to $18.75M actual against a $19.20M plan (a small but story-worthy -2.3% miss). Backs the nine-visual first draft and the five-visual rebuild on the Chapter 12 companion page.
Yuki’s session
The composite 1:47 a.m. reroll session behind Yuki’s dashboard-polish scenario in Chapter 13. Four iterations (none verified against any rubric) and five diagnostics (all firing) — the canonical 5-of-5 vibe-charting pattern the chapter defines, with contrast-ratio miss documented (3.1:1 shipped against WCAG’s 4.5:1 minimum). Backs Figure 13.2 on the Chapter 13 companion page.
Marcus’s compromise dashboard
The six-visual compromise Page 1 Marcus would build from a merged BRD (VP + regional director), alongside the two-BRD alternative: one small chart per audience, each with a named decision and a specified read-time budget. Locks the shape of the merged-BRD failure so the compromise dashboard and the two-BRD solution stay in lockstep with Chapter 2’s worked example. Backs Figure 2.4 on the Chapter 2 companion page.
Lina’s channel ROMI
The five-channel B2B marketing-attribution dataset behind Lina’s clustered-column failure and its Big-Idea rebuild. Four metrics per channel (leads, cost per lead, conversion rate, ROMI) so the reader can reproduce both the three-series clustered column and the single-encoding sorted bar of ROMI. Paid Social top ($3.40) and Display bottom ($0.80) are locked to keep the Big Idea’s exact wording. Backs Figure 3.5 on the Chapter 3 companion page.
Nadia’s Q3 revenue by region
The four-region CFO reporting snapshot behind Nadia’s Wednesday-morning comparison in Chapter 7. Quarter aggregate lands on plan ($28.4M); NA misses plan by -5.3% while APAC, EMEA, and LATAM all beat theirs — the on-plan-hides-miss shape the five-principle rubric was built to catch. Both actual and plan travel on every row so the reader can reproduce both the Copilot first-draft chart and the hand-built chart. Backs Figure 7.4 on the Chapter 7 companion page.
Kenji’s ticket-tier divergence
Twelve months of tier-2 ticket volume and mean resolution time behind Kenji’s Wednesday-mid-morning stacked-area failure and small-multiples rebuild. Nine months pre-automation, three months post-automation; pre-vs-post averages land on the ~19% volume drop / ~34% resolution-time climb the Big Idea cites. Backs Figure 4.4 on the Chapter 4 companion page.
Tomás’s governance dashboard
The fifteen-visual dense governance dashboard Tomás built on Friday afternoon and the keep/remove verdict for each visual after the decluttering pass. Every visual carries a specific reason; the kept set spans four chart families; the funding-decision bar chart is where the eye lands first in the rebuild. Backs Figure 5.3 on the Chapter 5 companion page.
Ch 5 practice — Q3 revenue vs plan
The four-region revenue snapshot behind Chapter 5’s Practice with us problem. The quarter lands within 2% of plan, which reads as fine; underneath it EMEA finished at 85% of plan, $1.2M short, while every other region beat theirs. The chart ships with every default decoration switched on, and the exercise is to inventory each element as encodes / helps read / neither before rebuilding. Practice-scope: no anchor character owns it. Held to the same $M scale as Nadia’s Q3 region data so the two revenue-by-region figures stay comparable.
Sara’s EMEA CSAR loop
The two-attempt CSAR-loop trace behind Sara’s Tuesday-morning one-page EMEA Q3 summary. Cold attempt (four pages, seven visuals, nobody trusts it) vs CSAR attempt (six clarifying questions, one-page ship in twelve minutes). Compression: 4p/7v → 1p/3v. Backs Figure 8.3 on the Chapter 8 companion page.
Diego’s campaign-attribution report
Diego’s Q3 campaign-attribution report: three cold-pass failure modes (wrong measure, wrong grain, wrong narrative) alongside the six Crystallize questions that surface them and the three Refine corrections. Reconciled against the finance close before ship. Backs Figure 9.2 on the Chapter 9 companion page.
Asha’s escalation drivers
Three AI-ranked escalation drivers with raw multipliers (age-over-7-days at 4.1x, Enterprise tier at 2.3x, Identity product line at 1.8x) and Asha’s confound-check verdict per driver: one clean survivor, one that survives when scoped to a feature area, one that collapses as a proxy for customer size. Backs Figure 10.5 on the Chapter 10 companion page.
Mei’s revenue-cycle Prep for AI
Five CFO-shaped questions asked against Mei’s revenue-cycle semantic model, before and after Prep for AI. Before: two worked, three broke in three different ways. After: four worked, one correctly refused (causal question). Each closed gap names the Prep component that fixed it (schema simplification, verified answers, AI instructions). Backs Figure 11.2 on the Chapter 11 companion page.
Anand’s four-moves capstone
Anand’s headcount-allocation dashboard through the four responsible-AI moves (Provenance, Fairness, Decision-Boundary, Disclosure). Fairness reveals the under-30 cohort at 22% attrition (2.75x aggregate); Decision-Boundary adds the escalation clause for downstream archival risk; two substantive issues surface that would have shipped invisible. Backs Figure 14.4 on the Chapter 14 companion page.
About the imperfections
Each dataset includes intentional data-quality issues — null values, naming inconsistencies, definitional drift, mid-window state changes. The imperfections are the point. A chapter exercise that pretends the data is perfect is teaching the wrong discipline.
Each dataset's SCHEMA.md documents every imperfection explicitly. The chapter exercises in the book reference the imperfections by name where they matter — the regional-naming reconciliation in Chapter 4, the cross-channel attribution reconciliation in Chapter 11, the self-assessment-versus-operational-signal triangulation in Chapter 13.
Reproducibility
The four cross-chapter datasets are generated by Node scripts in the companion-site repository at scripts/datasets/. Each script uses a seeded pseudo-random number generator (mulberry32 with a documented hex seed per script), so rerunning produces byte-identical output. The CSVs you download here are the same CSVs anyone else downloading at any time will receive.
Priya’s Q3 sales dataset is generated deterministically from a canonical JSON source in the book repository at applied-ai/ai-directed-viz/data/priya-q3-dashboard.json; the source SHA-256 travels with each derivative so drift between the book and this site is detectable. Aisha’s AIRS cohorts, Rohan’s services dashboard, Yuki’s session, Marcus’s compromise dashboard, Lina’s channel ROMI, and Nadia’s Q3 region datasets follow the same pattern — each has a canonical JSON source in applied-ai/ai-directed-viz/data/ and a dedicated generator script under applied-ai/ai-directed-viz/scripts/generate-*.mjs.
If you want to regenerate the datasets with a different seed for your own teaching purposes, the scripts are open in the companion-site repo.
License and reuse
Released for teaching use alongside The Defensible Decision. Use freely in classes, workshops, and your own practice. No attribution required beyond what's already in the dataset READMEs.