<!-- Generated by scripts/generate-llms-resources.mjs; do not edit. -->

# Starbucks AI-Driven BI Capstone

> An evidence-gated MBA capstone for framing, analyzing, and defending a Starbucks business intelligence recommendation with AI assistance.

Canonical URL: https://www.thedefensibledecision.com/starbucks/

## Overview

### Starbucks AI-Driven BI Capstone

#### Starbucks Operations Case

> **Independent educational case.** This project is not affiliated with, sponsored by, or endorsed by Starbucks Corporation. Starbucks is used nominatively to identify the company discussed. Maya Rodriguez and the engagement are fictional.

##### Status

- **Local capstone status: ready.** The case, BRD, four labs, templates, dashboard, Worked Solutions, rubric, and validation gates are complete for local course use.
- **Public companion status: ready.** This student route provides the approved capstone materials and hash-pinned course inputs.

##### Purpose of the Case

This case turns the disciplines in *The Defensible Decision: A Guide to AI-Assisted Business Analytics* into one connected BI assignment.

The book's premise is simple: principles first, AI later, judgment always. The Starbucks case puts that premise under pressure. You will work with observed public samples, simulated workflows, and missing operational evidence. Your job is to produce a useful recommendation without letting a polished dashboard, a convenient join, or an AI-generated explanation claim more than the data supports.

The executive decision is whether to fund **Centralized Analytics**, **Unrestricted Self-Service**, or **Governed Self-Service** for store health. The analytical challenge is to decide what each evidence source earns, what it cannot establish, and what the organization must measure next.

##### Overall Objectives

The case has four overall objectives:

1. **Frame the decision before opening the tool.** Define the audience, decision, question, comparison, time window, and success condition in a Business Requirements Document.
2. **Build an evidence chain rather than a collection of charts.** Move from operations workflow, to customer signal, to join boundary, to a bounded pilot.
3. **Direct and critique AI-assisted analysis.** Use AI to accelerate exploration and drafting while keeping metric choice, interpretation, and decision rights with the analyst and reviewer.
4. **Defend an executive recommendation.** State the preferred operating model, trade-offs, risks, missing measurements, refused claims, and evidence that would change the decision.

##### Learning Objectives

After completing the case, you will be able to:

1. Write a decision-ready BRD using the book's context before any chart discipline.
2. Turn a business question into a comparison, a defensible visual form, and a one-sentence takeaway.
3. Apply focus, hierarchy, and accessible labeling so evidence class, source, period, and denominator remain visible.
4. Distinguish observed evidence, simulated workflow findings, and missing data before making a recommendation.
5. Audit source relationships by grain, key, time, and selection mechanism before joining tables.
6. Use the CSAR loop to crystallize the question, scope the evidence, assemble a draft, and refine it through critique.
7. Detect misleading defaults, false causal language, unsupported AI narratives, and other reasons to ship, iterate, or discard an output.
8. Design a governed analytics model with certified measures, local investigation rights, Copilot guardrails, and consequential human review.
9. Write a provenance statement that names data sources, simulations, analytical choices, AI assistance, reviewers, and decision boundaries.
10. Present a defensible decision to a skeptical executive and state what evidence would justify revisiting it.

##### How the Book Guides the Case

| Book discipline | Case application |
| --- | --- |
| **Context before any chart** | Complete the BRD before opening Power BI or Copilot. |
| **Story from data to argument** | Give every visual one decision-relevant takeaway and carry one bounded finding per lab into the executive brief. |
| **Question-to-form matching** | Choose each chart from the analytical question and comparison, not from a default gallery. |
| **Focus and accessibility** | Remove decoration, preserve hierarchy, and keep labels readable across desktop and mobile. |
| **Principles first, AI later** | Understand the measure and evidence boundary before asking AI to explain or compose it. |
| **CSAR direction** | Crystallize, Scope, Assemble, and Refine AI-assisted work instead of accepting the first draft. |
| **Prep for AI** | Define stable keys, trustworthy measures, verified answers, and AI instructions before expecting reliable responses. |
| **Dashboard composition** | Lead with the decision, limit the executive surface, and move investigation detail behind the headline view. |
| **Failure-mode review** | Check for wrong measure, wrong grain, wrong narrative, misleading scale, and unsupported certainty; then ship, iterate, or discard. |
| **Responsible AI** | Preserve provenance, examine fairness risks, disclose AI assistance, and keep consequential decisions under human authority. |

##### Evidence Contract

| Evidence class | Valid use | Not valid for |
| --- | --- | --- |
| **Observed public data** | Bounded descriptions of the named sample, period, and selection mechanism | Starbucks-wide or causal conclusions |
| **Simulated data** | BI workflow practice and testable hypotheses | Claims about actual Starbucks customers or operations |
| **Missing evidence** | Defining the field, grain, key, denominator, or history needed next | Filling gaps with persuasive narrative |

Every headline finding must identify its evidence class. Every recommendation must state what the evidence earns and what it does not earn.

##### Choose Your Path

| Path | Start here | Purpose |
| --- | --- | --- |
| Understand the decision | [Case](index.html?area=overview&doc=case-plan) | Read the narrative, Big Idea, learning objectives, four-lab argument, and success criteria. |
| Audit provenance | [Sources & Licenses](index.html?area=overview&doc=source-notes) | Review source, license, sample, period, evidence class, and publication boundary. |
| Complete the assignment | [Student Path](index.html?area=student-path) | Work through the BRD, data package, Labs 01-04, synthesis, and blank executive brief. |

> **Student-work boundary:** Complete the Student Path independently. A defensible submission explains its evidence, limits, and recommendation without relying on an answer key.

##### Completion Standard

The case is complete when the executive can answer:

- What should I approve?
- Why does the evidence support it?
- What is observed, simulated, or still missing?
- Which comparison is useful but not causal?
- What may local teams and Copilot do, and what remains a human decision?
- Who acts next, what will they measure, and what result would change the recommendation?

The strongest submission must not disagree with its own evidence.


### The Missed Opportunity

#### The Missed Opportunity: When Customer Voice Cannot Reach Store Operations

> **Independent educational case.** This case was developed for *The Defensible Decision* companion materials. It is not affiliated with, sponsored by, or endorsed by Starbucks Corporation. Starbucks is used nominatively to identify the company discussed. Maya Rodriguez and the engagement described below are fictional.

##### Capstone Status

- **Local capstone status: ready.** Teams may complete the full project from the local case package.
- **Public companion status: ready.** This student route provides the approved capstone materials and hash-pinned course inputs.

##### Big Idea

> Public Starbucks data cannot prove why operations changed. It can show analysts where evidence stops, which joins are missing, and how those limits should change an executive recommendation.

A polished dashboard can make limited evidence look complete. This case asks you to do the opposite: preserve uncertainty, refuse invalid joins, and let the evidence boundary change the recommendation.

##### The Case In One Sentence

Maya Rodriguez must choose an analytics operating model that gives store teams useful customer insight without letting speed, self-service, or AI outrun the evidence.

##### The Decision Situation

You are part of a BI team advising Maya, a fictional VP of Store Operations Analytics. Leadership is deciding what to fund next: a centralized reporting function, unrestricted local self-service, or a governed model that combines certified measures with local investigation.

Maya cannot wait for a perfect enterprise dataset. She must recommend a direction using the evidence available now while naming the measurements required before stronger operational decisions are safe.

The team must choose among three operating models:

1. **Centralized analytics:** corporate specialists define and distribute every metric. Definitions stay consistent, but local questions move slowly.
2. **Unrestricted self-service:** store and district teams build their own measures and reports. Local questions move quickly, but definitions and interpretations can fragment.
3. **Governed self-service:** certified measures support local investigation, with explicit evidence labels, defined authority, Copilot guardrails, and human review.

The wrong choice has two costs. Too much central control leaves store teams waiting for answers. Too little control lets teams act on conflicting metrics, selected samples, or AI-generated explanations that the data does not support.

Your task is not to prove that customer feedback predicted a corporate crisis. The available public data cannot establish that claim. Your task is to decide what the evidence earns, what remains hypothetical, and what the organization must measure next.

##### The Business Problem

Large organizations often separate customer experience, operations, marketing, and workforce data by function. Each team may have a defensible view of its own measures while no one owns the combined question: **is this store healthy, and what should a manager do next?**

For this case, assume leadership wants a Store Health experience that helps a manager answer:

- What changed?
- Where is the change concentrated?
- What evidence supports the finding?
- What action is within my authority?
- What is still unknown?

The difficulty is not a shortage of charts. It is the absence of a trusted connection between measures collected at different grains, in different periods, and for different purposes.

##### Learning Objectives

After completing the case, you will be able to:

1. **Classify evidence** as observed, simulated, or missing and state the decision use each class earns.
2. **Build bounded comparisons** that keep source, period, sample size, and denominator visible.
3. **Audit analytical relationships** by grain, key, time, and selection mechanism before joining sources.
4. **Distinguish description from explanation** and refuse causal or network-wide claims unsupported by the data.
5. **Translate patterns into tests** by specifying pilot candidates, holdouts, owners, and exit criteria instead of recommending immediate deployment.
6. **Design a governed analytics model** that separates certified measures, local investigation, Copilot assistance, and consequential human decisions.
7. **Defend an executive recommendation** with trade-offs, risks, missing measurements, refused claims, and evidence that would change the decision.

##### Capstone Project Definition

This is the culminating project for the AI-assisted business intelligence course. A capstone must integrate the course's disciplines in one authentic decision, not repeat four disconnected lab exercises.

Your team will frame the business decision, prepare and audit evidence, build a focused analytical experience, direct and critique AI-assisted work, design governance boundaries, and defend one recommendation to an executive audience.

| Project element | Capstone contract |
| --- | --- |
| **Format** | Team submission; Teams of 3-4 unless the instructor approves another format |
| **Role** | BI advisory team to Maya Rodriguez, fictional VP of Store Operations Analytics |
| **Decision** | Choose and defend an analytics operating model for Store Health |
| **Primary audience** | Executive sponsor, store operations leadership, and district leadership |
| **Working tools** | Power BI or an equivalent BI tool, with disclosed Copilot or other instructor-approved AI assistance |
| **Evidence** | Public observed samples, explicitly simulated workflows, and named missing measurements |
| **Final outcome** | An executive-ready recommendation supported by an auditable evidence chain and a bounded implementation roadmap |

The capstone is not complete because the dashboard renders. It is complete when the team can defend the decision, the evidence behind it, the limits of that evidence, and the authority boundaries around AI-assisted action.

##### Course Competencies Demonstrated

| Course competency | Evidence expected in the capstone |
| --- | --- |
| Context before any chart | An accepted BRD naming audience, decision, question, comparison, time window, and success condition |
| Story and chart selection | One-sentence takeaways and visual forms matched to the required comparison |
| Focus, hierarchy, and accessibility | A restrained executive surface with readable labels, evidence context, and intentional emphasis |
| Evidence and data readiness | Source, grain, period, denominator, selection mechanism, key audit, and missing-data specification |
| CSAR and AI direction | A traceable question-to-draft-to-critique-to-refinement cycle for representative AI-assisted work |
| AI output verification | Human-checked calculations, claims, visual scales, and narrative conclusions |
| Dashboard composition | Clear decision hierarchy, bounded interactivity, and investigation detail separated from executive summary |
| Failure-mode review | Explicit ship, iterate, or discard decisions for misleading or unsupported outputs |
| Responsible AI and governance | Provenance, fairness considerations, disclosure, decision boundaries, and human approval points |
| Executive decision quality | A recommendation with trade-offs, risks, owners, exit tests, refused claims, and revisit evidence |

##### Team Scope and Roles

Teams of 3-4 should assign clear primary ownership while preserving shared accountability for the final decision.

| Role | Primary responsibility |
| --- | --- |
| **BI lead** | Maintain the decision frame, coordinate milestones, and reconcile the final recommendation |
| **Evidence and source auditor** | Validate provenance, grain, periods, denominators, joins, and refused claims |
| **Report and dashboard designer** | Build the model, measures, visuals, interactions, hierarchy, and accessibility checks |
| **AI and governance reviewer** | Maintain the AI-use record, challenge generated output, define guardrails, and coordinate executive defense |

A three-person team may combine the final two roles. Every team member must be able to explain the evidence contract, the recommendation, and one reason the recommendation could be wrong. Submit a brief contribution statement identifying primary ownership and cross-review performed by each member.

##### Evidence Contract

The case deliberately combines three evidence classes. They are not interchangeable.

| Evidence class | What it can support | What it cannot support |
| --- | --- | --- |
| Observed public data | Bounded descriptions of the named sample, period, and selection mechanism | Network-wide or causal claims about Starbucks |
| Simulated data | BI workflow practice and testable hypotheses | Claims about actual Starbucks customers or operations |
| Missing evidence | A specification for the field, key, grain, denominator, or history needed next | A persuasive story that fills the gap |

Every recommendation must identify its evidence class.

##### The Four-Lab Argument

The labs are one cumulative argument. Each corrects a tempting analytical shortcut and produces an artifact for the executive brief.

###### 1. Store Health

**Big Idea:** An operations dashboard can optimize the variables it contains while missing a larger change visible in customer voice.

Use a simulated ordering extract to build an operations dashboard, then compare its visible differences with an observed two-store customer-review sample. Do not join the sources. Identify the store-by-time measurement missing between them.

###### 2. Store Experience

**Big Idea:** The strategic difference between two stores can live in rating shape, coverage, and sources of delight rather than in the average alone.

Compare Broadway and Reserve Roastery using rating shape, review coverage, and a consistent language-treatment rule. Decide which experience ideas might be transferable and which conclusions remain specific to two NYC stores.

###### 3. Evidence Boundary

**Big Idea:** Geographic overlap can tell analysts where to investigate; it cannot make customer reviews a proxy for employee conditions.

Compare customer-review geography, a 2021 store snapshot, and public workforce-event records at state level. Build a queue for primary research without implying that the sources refer to the same stores or establish a labor-and-customer causal relationship.

###### 4. Offer Completion

**Big Idea:** Completion, not views, is the action metric; segment differences define a pilot hypothesis, not a production targeting rule.

Use a simulated campaign extract to separate offer attention from completion. Identify segment-offer cells that justify a randomized pilot with a holdout, not a production targeting rule.

##### Evidence-to-Decision Map

| Stage | Learning objective demonstrated | Artifact produced | Contribution to the final decision |
| --- | --- | --- | --- |
| BRD | Frame the audience, decision, question, comparison, and success condition | Accepted team framing | Defines what the analysis must change |
| Lab 01 | Classify evidence and expose missing measurement | Operations view, customer-voice view, reality gap | Establishes why Store Health needs a customer-voice layer |
| Lab 02 | Build bounded comparisons and preserve coverage | Rating-shape, period, and theme comparison | Separates transferable experience ideas from flagship-specific conclusions |
| Lab 03 | Audit grain and reject a false join | Source audit, state views, research queue, store-month specification | Converts weak overlap into a primary-research investment |
| Lab 04 | Translate a pattern into a fair test | Receipt funnel, offer comparison, candidate cells, holdout design | Defines a bounded pilot instead of direct targeting |
| Executive brief | Defend a governed decision | Recommendation, trade-offs, roadmap, risks, and refused claims | Requests a specific approval and names the next owner |

##### Milestones and Decision Gates

Each gate is part of the capstone assessment. Do not move forward with an unresolved evidence or decision defect.

| Milestone | Required evidence | Exit gate |
| --- | --- | --- |
| **1. Team charter and BRD** | Team roles, decision owner, users, question, comparison, success condition, and scope boundaries | The team can explain the decision without opening a dataset |
| **2. Operations and customer signal** | Completed Labs 01-02, evidence labels, sample sizes, coverage limits, and reality gap | The team can state what the dashboard sees and what it cannot explain |
| **3. Boundary and pilot** | Completed Labs 03-04, grain audit, research queue, receipt funnel, and holdout design | The team refuses false joins and converts patterns into testable actions |
| **4. Synthesis and design review** | One finding per lab, operating-model choice, dashboard draft, AI-use record, and provenance statement | Every headline claim traces to a source and survives adversarial review |
| **5. Executive defense** | Final submission package and rehearsal against the rubric | A skeptical executive can identify the approval, trade-off, risk, next owner, and revisit evidence |

##### What a Store-Level Dashboard Might Need

The conceptual mock-up below illustrates the decision experience Maya wants. Its store, values, alerts, operating measures, and recommendation are fictional. The public review data in this case cannot populate the operational or employee measures shown.

![Conceptual Store-Level Voice-of-the-Customer dashboard with fictional values](assets/store-voc-dashboard.png)

The worked solution includes an evidence-bounded dashboard built only from the observed two-store review sample. It shows rating shape, review coverage, annual patterns, and transparent keyword-theme mentions. It deliberately excludes wait time, staffing, transactions, employee measures, correlations, and causal recommendations.

##### Decision To Make

Recommend one analytics operating model and answer:

- Which measures should be certified centrally?
- What should store and district teams be allowed to investigate locally?
- Where may Copilot assist, and where is human review mandatory?
- What store-month data model is required before stronger operational claims are possible?
- What should the first bounded pilot test?

##### Capstone Submission Package

Submit one coherent package containing:

1. **Team charter and contribution statement:** named roles, primary ownership, and cross-review responsibility.
2. **Final BRD:** the accepted decision contract and any documented changes made after evidence review.
3. **BI report or equivalent interactive artifact:** a focused executive surface plus the required Store Health, Customer Voice, Evidence Boundary, and Campaign Pilot analysis.
4. **Evidence audit:** source, grain, period, denominator, selection mechanism, join decision, and missing store-month specification.
5. **Capstone Journal:** working notes, completed tasks and methods, Copilot prompts, representative CSAR artifacts, human critiques, provenance, verification, and final decision owners.
6. **One-page executive brief:** recommendation, evidence snapshot, trade-offs, roadmap, risks, refused claims, and decision request.
7. **Executive presentation:** the visuals and speaking plan needed for the timed defense below.

The four lab outputs are working evidence. The submission should synthesize them rather than paste them together unchanged.

##### Journal, Copilot Use, and Academic Integrity

AI assistance is part of the capstone workflow. It does not transfer accountability away from the team.

| Allowed | Required | Not permitted |
| --- | --- | --- |
| Brainstorm questions, draft formulas, propose visual alternatives, summarize bounded findings, and critique prose | Disclose tools and models; preserve representative prompts, outputs, critiques, and refinements; verify calculations and sources; identify the human decision owner | Treat AI output as evidence, fabricate sources or values, conceal material AI assistance, use unsupported causal language, or submit answer-key reasoning as team work |

**AI output is not evidence.** The underlying data, documented calculation, and verified source are evidence. Consequential recommendations involving staffing, customer targeting, workforce interpretation, or automated action require explicit human review.

> **Independent-work boundary:** Complete the team submission independently. Reusing an answer key's charts, prose, calculations, or recommendation does not demonstrate capstone competence.

##### Executive Presentation

Unless the instructor specifies another format, deliver a **7-minute executive briefing** followed by questions.

The briefing must include:

- one decision request and one recommended operating model;
- two material trade-offs;
- three evidence-backed findings spanning observed, simulated, and missing evidence;
- one visual reality gap or evidence-boundary finding;
- one bounded pilot or implementation roadmap;
- one refused claim;
- one next action with an owner and exit test.

Use the report as evidence, not as a script. Every team member should be prepared to answer questions about source quality, AI use, and what would change the recommendation.

##### Evaluation Rubric

| Dimension | Weight | Strong performance |
| --- | ---: | --- |
| **Evidence quality and refusal discipline** | 30% | Claims trace to source, grain, period, denominator, and selection mechanism; false joins and unsupported conclusions are explicitly refused |
| **Decision framing and analytical argument** | 20% | BRD, Big Idea, comparisons, findings, and recommendation form one coherent decision chain |
| **Analysis and visual design** | 15% | Measures are correct; chart forms match questions; hierarchy, focus, accessibility, and interactivity support the audience |
| **AI direction, provenance, and responsible use** | 15% | AI work is disclosed, critiqued, verified, and bounded by clear human decision rights |
| **Operating model, pilot, and roadmap** | 10% | Governance, local authority, missing measurement, owners, risks, and exit tests are concrete and feasible |
| **Executive communication and team defense** | 10% | The briefing is concise, the trade-off is explicit, all members demonstrate ownership, and answers remain inside the evidence boundary |

##### Capstone Completion Gate

A submission passes the capstone gate when a skeptical executive can answer:

- What should I approve, and why?
- Which findings are observed, simulated, or still missing?
- Which comparison is useful but not causal?
- Where does the current evidence stop?
- What should local teams be allowed to investigate?
- Which AI-assisted actions require human review?
- What new measurement or pilot would justify revisiting the decision?

The recommendation may differ from the Worked Solution. It earns a strong evaluation when the evidence chain is complete, the trade-off is explicit, and the refused claims remain visible.

##### Claims The Case Refuses

- Public customer reviews prove why Starbucks performance changed.
- Simulated channel or campaign differences describe actual Starbucks behavior.
- Shared state geography establishes same-store overlap.
- Customer demographics cause offer completion or justify exclusion.
- A polished dashboard can substitute for missing operational history.

##### Start The Case

1. Read the [Business Requirements Document](../student-path/brd.md) and accept or challenge its decision contract.
2. Open the [Student Path](../student-path/README.md) and complete the four labs in order.
3. Carry one bounded finding from each lab into the synthesis table.
4. Complete the blank executive brief and test it against the learning objectives above.
5. Present the recommendation, its strongest uncertainty, and the evidence that would change it.

##### Continue

- Return to [Overview](../index.html?area=overview&doc=start-here) for objectives and book alignment.
- Review [Sources & Licenses](../index.html?area=overview&doc=source-notes) before distributing data or presenting source claims.
- Open the [Student Path](../index.html?area=student-path) to begin the capstone.

##### Source And Trademark Note

The [Source, License, And Evidence Notes](../index.html?area=overview&doc=source-notes) identify every observed and simulated source, recorded license, sample, period, and public-redistribution boundary. All Starbucks trademarks and brand names remain the property of their respective owners. No private Starbucks data is used.


### Business and Market Context

#### Business and Market Context

Use this brief to place observed review patterns inside the business environment in which they occurred. The events below are **context variables**, not explanatory variables in the provided datasets.

> **Causal boundary:** A date alignment does not prove that an event changed ratings, review volume, customer behavior, labor conditions, or store performance. The case lacks verified event exposure, a counterfactual, and a same-store operating history. Treat each event as a reason to ask a better question, not as an answer.

##### How to Use Context Without Claiming Causation

For every event that overlaps a chart:

1. State what happened and cite the source.
2. Check whether the event overlaps the data period and the named store.
3. Inspect rating, review-volume, and theme patterns before and after the date.
4. Name at least one competing explanation.
5. Specify the missing measure needed to test the event's effect.
6. Use language such as **coincides with**, **provides context for**, or **motivates a test**.

Do not write **caused**, **drove**, **led to**, or **resulted in** unless a separate design establishes that relationship.

###### Annual Sample-Size Rule

Use a minimum annual review count of $N=10$ before treating a store-year mean as a reliable trend point. Keep years below that threshold visible as hollow thin-sample observations. A dashed connector may join adjacent observed years when at least one endpoint is thin, but it is directional only; never use a solid connector for that segment or bridge a missing year. A missing year is a gap, not a value to interpolate.

In the provided data, thin store-years are:

- Broadway: 2008 ($N=2$), 2010 ($N=5$), 2022 ($N=5$), 2023 ($N=5$), and 2024 ($N=1$);
- Reserve Roastery: 2018 ($N=4$) and 2024 ($N=1$).

The threshold is a display and interpretation safeguard, not a claim that $N=10$ makes a convenience sample representative. Selection bias, platform effects, language, and changing visitor mix still apply.

##### Event Timeline

| Marker | Date | Event | Why it matters for analysis | What the case can and cannot do |
| ---: | --- | --- | --- | --- |
| — | 2008-07-01 | Starbucks announced plans to close 600 underperforming U.S. stores and cut up to 12,000 positions during an economic downturn; Howard Schultz had returned as CEO in January. | The Broadway review history begins during a corporate retrenchment shaped by weak consumer spending, competition, and prior overexpansion. | Use as an opening market regime, not a store-specific treatment. The review sample does not identify closure exposure or a stable pre-event baseline. |
| — | 2011-01-19 | Starbucks launched mobile payment across U.S. company-operated stores. | Payment digitization changed convenience and app engagement before order-ahead became available. | The reviews do not record payment method or app adoption, so they cannot estimate a mobile-payment effect. |
| — | 2015-03-18 | The Race Together campaign drew social-media backlash after Starbucks encouraged conversations about race in stores. | Brand activism and customer outcry can alter reputation, review language, and who chooses to post. | Treat as national reputation context. The two-store sample cannot identify campaign exposure or isolate its effect. |
| — | 2015-09-22 | Starbucks rolled out Mobile Order & Pay nationwide. | Order-ahead can change queueing, handoff, customization, throughput, and the relationship between digital convenience and in-store experience. | The review files contain no order channel, wait time, handoff accuracy, or adoption field; causal claims require those measures. |
| — | 2017-04-03 | Kevin Johnson succeeded Howard Schultz as CEO. | The succession marked a leadership and technology-strategy transition before the Philadelphia incident and Roastery opening. | Use as company context only. Annual reviews cannot attribute store experience to executive leadership. |
| — | 2018-04-12 | Two Black men were arrested at a Philadelphia Starbucks; the incident went viral and prompted protests and accusations of racial profiling. | Customer trust, inclusion, and brand reputation can change review content and who chooses to post. | Use as national brand context. The two NYC review files do not identify event exposure or establish an effect. |
| 1 | 2018-12-14 | The Starbucks Reserve Roastery New York opened in Chelsea. | This is the opening boundary for the Roastery sample and explains why its review history begins in late 2018. | Use to define the sample window. Do not interpret opening-year enthusiasm as a stable baseline without checking volume and selection. |
| — | 2019-05-06 | Microsoft and Starbucks described Azure-hosted reinforcement learning, connected equipment, predictive maintenance, and digital traceability programs. | The program establishes a technology-capability baseline for personalization and store operations. | Capability does not prove deployment at either sampled store, adoption by customers, or an effect on ratings. |
| 2 | 2020-03-20 | Starbucks closed U.S. company-operated cafes and shifted service toward drive-thru and delivery during the COVID-19 pandemic. | Store access, visit purpose, service channels, staffing, and review-generation behavior changed simultaneously. | Treat 2020 as a structural disruption. Annual reviews cannot isolate a pandemic effect from channel, selection, or composition changes. |
| 3 | 2021-12-09 | Workers at one Buffalo store voted to join Workers United, the first unionized U.S. company-owned Starbucks location at that time. | Labor relations became a visible strategic and reputational issue, potentially changing employee and customer narratives. | Use as national labor context. A6 and the review files lack a verified same-store key, so no labor-to-rating join is justified. |
| 4 | 2022-04-04 | Kevin Johnson retired and Howard Schultz became interim CEO. | Leadership transition can coincide with changes in strategy, investment, communications, and operating priorities. | Annotate the period, but do not attribute a store rating change to executive leadership without intervening measures and a comparison design. |
| — | 2022-09-13 | Starbucks announced a Reinvention Plan that included major investment in North American store equipment and workflows. | The plan directly raises hypotheses about throughput, customization, equipment, staffing, and partner experience. | The review sample has no implementation date, store exposure, equipment, staffing, or throughput measures. It cannot evaluate the plan's effect. |
| 5 | 2023-03-20 | Laxman Narasimhan assumed the CEO role. | A second leadership transition occurred during the observed Roastery decline. | Treat leadership as a competing context hypothesis, not a causal explanation. |
| 6 | 2023-12-20 | Starbucks faced boycott calls, protests, and reported store vandalism amid controversy over representations of its position on the Israel-Gaza war. | Customer outcry and social-media attention may affect visits, review volume, sentiment, and the mix of people posting. | The event occurred late in 2023. The annual sample cannot separate its effect from other 2023 conditions. |
| — | 2024-04-30 | Starbucks described fiscal Q2 2024 as a difficult quarter that did not meet expectations. | This is a later business outcome that helps frame why leadership and turnaround decisions followed. | The quarter ended after the review sample. It must not be used as proof that the plotted reviews predicted company performance. |
| — | 2024-09-09 | Brian Niccol began as chairman and CEO after Starbucks announced the appointment on August 13. | The transition is important for forward-looking strategy and any extension of the case. | It occurred after the review sample ended on 2024-01-08. Do not use it to explain any plotted rating or review count. |

##### Analysis Prompts

Use the timeline to enrich, not replace, the evidence analysis:

- **Coverage:** Does review volume change around an event? Could access or posting behavior explain the change?
- **Composition:** Did the mix of stores, ratings, languages, or themes change?
- **Timing:** Is the event early, late, or outside the observed year? Annual aggregation can hide sequence.
- **Alternative hypotheses:** Could seasonality, platform changes, tourism, inflation, staffing, menu changes, or sample selection explain the same pattern?
- **Missing measures:** What store-month fields would distinguish the hypotheses: transactions, channel mix, wait time, staffing, complaints, local event exposure, or media intensity?
- **Decision relevance:** Would the context change the recommendation, the confidence level, or only the next measurement request?

##### Required Context Note

For each time-series visual in the team submission, include a note in this form:

> **Context marker [number]:** [event] occurred during this period. The chart shows [bounded pattern]. It does not prove the event caused the pattern because [missing exposure, comparison, or measure]. We would test the relationship using [specific design and fields].

##### Source Ledger

| Event | Source |
| --- | --- |
| 2008 store restructuring | Reuters, [“Starbucks to cut up to 12,000 jobs, close 600 stores,” July 1, 2008](https://www.reuters.com/article/world/starbucks-to-cut-up-to-12000-jobs-close-600-stores-idUSN01299942/) |
| Nationwide mobile payment | Nation’s Restaurant News, [“Starbucks rolls out mobile payment,” January 2011](https://www.nrn.com/restaurant-technology/starbucks-rolls-out-mobile-payment) |
| Race Together backlash | Reuters, [“Starbucks brews up backlash with debate on U.S. race relations,” March 18, 2015](https://www.reuters.com/article/business/starbucks-brews-up-backlash-with-debate-on-us-race-relations-idUSL2N0WK1PU/) |
| Mobile Order & Pay | Eater, [“Starbucks Rolls Out Mobile Order and Pay Ahead System Nationwide,” September 22, 2015](https://www.eater.com/2015/9/22/9369483/starbucks-rolls-out-mobile-order-pay-ahead-system-nationwide-app-ban-lines) |
| Johnson CEO | Starbucks Investor Relations, [“Starbucks Announces New Leadership Structure to Drive Next Wave of Global Growth,” December 1, 2016](https://investor.starbucks.com/news/financial-releases/news-details/2016/Starbucks-Announces-New-Leadership-Structure-to-Drive-Next-Wave-of-Global-Growth/default.aspx) |
| Philadelphia arrests and protests | Reuters, [“Arrest of two black men in Starbucks sparks protests,” April 17, 2018](https://www.reuters.com/news/picture/arrest-of-two-black-men-in-starbucks-spa-idUKRTX5RLEL/) |
| New York Roastery opening | Daily Coffee News, [“Starbucks Opens 23,000-Square-Foot New York Reserve Roastery,” December 13, 2018](https://dailycoffeenews.com/2018/12/13/starbucks-opens-23000-square-foot-new-york-reserve-roastery/) |
| Microsoft digital program | Microsoft Source, [“Starbucks turns to technology to brew up a more personal connection with its customers,” May 6, 2019](https://news.microsoft.com/source/features/digital-transformation/starbucks-turns-to-technology-to-brew-up-a-more-personal-connection-with-its-customers/) |
| COVID-19 operating shift | NPR, [“Starbucks Responds to COVID-19,” March 20, 2020](https://www.npr.org/sections/coronavirus-live-updates/2020/03/20/819359246/make-that-quad-long-shot-grande-in-a-venti-cup-to-go-starbucks-responds-to-covid) |
| Buffalo union vote | Reuters, [“Starbucks workers vote to unionize at Buffalo, New York, store,” December 9, 2021](https://www.reuters.com/business/starbucks-loses-bid-stave-off-labor-union-buffalo-new-york-2021-12-09/) |
| Schultz interim CEO | Starbucks Corporation release mirrored by Nasdaq, [“Starbucks Announces Leadership Transition,” March 16, 2022](https://www.nasdaq.com/press-release/starbucks-announces-leadership-transition-2022-03-16) |
| Reinvention Plan | CNBC, [“Starbucks hikes forecast, outlines plans for automated stores and more rewards benefits,” September 13, 2022](https://www.cnbc.com/2022/09/13/starbucks-projects-long-term-earnings-revenue-growth-in-double-digits-as-it-implements-new-strategy.html) |
| Narasimhan CEO | Starbucks release mirrored by Nasdaq, [“Laxman Narasimhan Assumes Role of Starbucks Chief Executive Officer,” March 20, 2023](https://www.nasdaq.com/press-release/laxman-narasimhan-assumes-role-of-starbucks-chief-executive-officer-2023-03-20) |
| Boycott calls and protests | BBC, [“Starbucks blames ‘misrepresentation’ after Israel Gaza protests,” December 20, 2023](https://www.bbc.com/news/business-67777506) |
| Fiscal Q2 2024 underperformance | Starbucks, [“Starbucks Reports Q2 Fiscal 2024 Results,” April 30, 2024](https://about.starbucks.com/press/2024/starbucks-reports-q2-fiscal-2024-results/) |
| Niccol CEO | Starbucks, [“Starbucks names Brian Niccol as Chairman and Chief Executive Officer,” August 13, 2024](https://about.starbucks.com/press/2024/starbucks-names-brian-niccol-as-chairman-and-chief-executive-officer/) |

##### Update Rule

This is a dated analytical context brief, not a live news feed. Add an event only when it is decision-relevant, sourced, date-specific, and accompanied by an explicit causal boundary. Recheck links and event status before reusing the case in a later course term.


### Source, License, and Evidence Notes

#### Source, License, And Evidence Notes

> **Independent educational case.** This material is not affiliated with, sponsored by, or endorsed by Starbucks Corporation. Starbucks is used nominatively to identify the company discussed. No private Starbucks data is used.

##### Use Status

- **Local capstone status: ready.** The local repository contains the hash-pinned inputs required for instruction and analysis.
- **Public companion status: ready.** This student route provides the approved capstone materials and hash-pinned course inputs.

##### Evidence Classes

- **Observed public data** supports bounded descriptions of the named sample, period, and selection mechanism.
- **SIMULATED data** demonstrates an analytical workflow or defines a hypothesis to test. It does not describe actual Starbucks behavior.
- **Missing evidence** identifies the field, grain, key, denominator, or time history required before stronger claims are possible.

##### Dataset Register

| ID | Source | Evidence class | Recorded license | Public companion status |
| --- | --- | --- | --- | --- |
| **A1** | [Starbucks Reviews Dataset](https://www.kaggle.com/datasets/harshalhonde/starbucks-reviews-dataset), contributed by `harshalhonde` | Observed, complaint-heavy selected sample | CC BY-NC 4.0 | [reviews_data.csv](../student-path/data/A1/reviews_data.csv) is cached on this unlisted route. It includes reviewer names, locations, and review text. |
| **A2** | Starbucks NYC Reviews, contributed by `muqaddasejaz`; upstream Kaggle source withdrawn after the course copy was captured | Observed two-store review sample | Open Data Commons PDDL 1.0 recorded for the database | [starbucks_ny_broadway.csv](../student-path/data/A2/starbucks_ny_broadway.csv) and [starbucks_ny_reserve_roastery.csv](../student-path/data/A2/starbucks_ny_reserve_roastery.csv) are cached on this unlisted route. |
| **A3** | [Starbucks Locations Worldwide 2021](https://www.kaggle.com/datasets/kukuroo3/starbucks-locations-worldwide-2021-version), contributed by `kukuroo3` | Observed historical store snapshot | CC0 1.0 | Approved for public companion distribution with source and snapshot-year disclosure. |
| **A6** | [Starbucks Union Election Results](https://www.kaggle.com/datasets/brandonconrady/starbucks-union-election-results), contributed by `brandonconrady` | Observed selected workforce-event records | CC0 1.0 | Approved for public companion distribution with source, period, and coverage disclosure. |
| **B1** | [Starbucks Customer Data](https://www.kaggle.com/datasets/ihormuliar/starbucks-customer-data), contributed by `ihormuliar` | **SIMULATED** Udacity campaign data | CDLA-Permissive 1.0 | Approved for public companion distribution only with a visible **SIMULATED** label. |
| **B3** | [Starbucks Customer Ordering Patterns](https://www.kaggle.com/datasets/likithagedipudi/starbucks-customer-ordering-patterns), contributed by `likithagedipudi` | **SIMULATED** ordering data | CC0 1.0 | Approved for public companion distribution only with a visible **SIMULATED** label. |

##### Publication Boundary

The local course repository and this public student route retain exact hash-pinned inputs so analyses can be reproduced:

- A1, A2, A3, A6, B1, and B3 are cached as immutable course inputs;
- the observed-review dashboard may export its derived date, store, rating, and theme-flag payload because it contains no raw review text;
- the internal case master plan, technical appendix, report contracts, and historical archive are excluded.

##### Analytical Boundaries

- A1 geography belongs to the reviewer record, not a verified Starbucks store.
- A2 covers two named NYC stores with unequal and nonparallel review histories.
- A3 is a 2021 store snapshot, not a current network census.
- A6 lacks A3's stable store key and cannot be joined deterministically to customer reviews.
- B1 and B3 are simulations regardless of their Starbucks branding.
- None of the sources forms the store-month operational history needed to explain changes in customer experience.

##### Trademark Note

Starbucks and the Siren logo are trademarks of Starbucks Corporation. The case stores the Siren SVG locally and uses it only to identify the company discussed. The mark remains the property of Starbucks Corporation; its presence does not imply corporate participation, sponsorship, or endorsement.

Kaggle and its wordmark are trademarks of Google LLC. The Sources & Licenses hero stores the official Kaggle site wordmark locally and uses it only to identify the source platform. Its presence does not imply Kaggle or Google sponsorship or endorsement.

##### Continue

- Return to [Overview](../index.html?area=overview&doc=start-here) for objectives and reader paths.
- Read the [Capstone Case](../index.html?area=overview&doc=case-plan) for the decision, milestones, submission package, and rubric.
- Open the [Student Path](../index.html?area=student-path) to complete the local capstone.


## Student Path

### Student Path: From Evidence to Executive Decision

#### Student Path: From Evidence to Executive Decision

##### Case In One Sentence

Can a governed analytics operating model help store teams act on customer voice without turning limited public evidence or simulated workflows into claims they cannot support?

You are part of a BI team advising Maya Rodriguez, a fictional VP of Store Operations Analytics. Maya must recommend how customer, operational, marketing, and workforce signals should reach store and district managers. The choice is between **Centralized Analytics**, **Unrestricted Self-Service**, and **Governed Self-Service**.

This is not a case about proving that customer feedback predicted a corporate crisis. The available public data cannot establish that claim. It is a case about deciding what the evidence supports, what remains hypothetical, and what the organization must measure next.

Read [The Missed Opportunity](../index.html?area=overview&doc=case-plan) before opening the BRD.

##### Your Assignment

**Audience:** Student BI team. \
**Decision owner:** Maya Rodriguez, fictional VP of Store Operations Analytics. \
**Decision:** Recommend an analytics operating model that supports local action without creating metric chaos. \
**Format:** Teams of 3-4 unless the instructor approves another format. \
**Final output:** One Capstone Submission Package containing the BI artifact, evidence audit, executive brief, completed Capstone Journal, and presentation.

**Big Idea:** Move from understanding the decision, to testing one evidence boundary at a time, to defending a recommendation that remains honest about what the data cannot prove.

##### The Sequence

```text
1. Form the team and accept the BRD
   ↓
2. Complete Labs 01-02: operations and customer signal
   ↓
3. Complete Labs 03-04: evidence boundary and pilot
   ↓
4. Synthesize and pass design review
   ↓
5. Defend the executive recommendation
```

Do not start by designing the final dashboard. The sequence is designed to prevent premature storytelling, false joins, and unsupported recommendations.

##### Why The Lab Order Matters

The labs are cumulative, not independent assignments:

1. **Lab 01 establishes the gap.** A simulated operations dashboard can rank the variables it contains, while observed customer voice reveals a signal it cannot explain.
2. **Lab 02 deepens the observed signal.** Rating shape, time coverage, and language treatment show why a mean alone is not enough.
3. **Lab 03 tests the temptation to overjoin.** Shared geography can prioritize research, but it cannot manufacture a verified store-level relationship.
4. **Lab 04 converts a pattern into a fair test.** Simulated campaign differences earn a randomized pilot with a holdout, not a production targeting rule.

Together, the labs move from **dashboard capability**, to **customer signal**, to **evidence boundary**, to **testable action**. Each leaves an artifact used in the final brief.

> **Student-work boundary:** Complete the case from this Student Path. Worked Solutions are a separate answer-key package for review after submission. Do not copy its charts, language, or recommendation into the team deliverable.

##### Student Packet

Use this page as the single starting point. The linked documents are the complete
working packet for the case:

Before starting the numbered workflow, read [Business and Market Context](../index.html?area=overview&doc=business-context)
in Overview for sourced leadership, customer, labor, and market events with explicit causal boundaries.

| Order | Document | Use |
| ---: | --- | --- |
| 1 | [BRD](brd.md) | Business Requirements Document: scope, deliverables, and acceptance criteria |
| 2 | [Dataset Package](dataset-package.md) | Review and download the public-source or simulated extracts assigned to each lab |
| 3 | [Lab 01](labs/01-store-health-dashboard.md) | Simulated operations dashboard and customer-voice reality check |
| 4 | [Lab 02](labs/02-store-comparison.md) | Two-store rating shape, themes, and language coverage |
| 5 | [Lab 03](labs/03-geographic-evidence-boundary.md) | State coverage, failed store join, and primary-research queue |
| 6 | [Lab 04](labs/04-offer-completion.md) | Simulated campaign funnel and randomized pilot design |
| 7 | [Executive Brief Template](reports/executive-brief.md) | Blank target structure for the team's final executive recommendation |
| 8 | [Capstone Journal](reports/journal.md) | Working notes, completed tasks, Copilot prompts, source decisions, verification, and team contributions |

The BRD provides the delivery contract. The Dataset Package provides the files.
The Overview's Business and Market Context brief supplies sourced events and
explicit causal boundaries. The labs provide the evidence and methods. The Executive
Brief and Capstone Journal are the target records for the final submission.

##### Capstone Artifact Manifest

Use this table as the team-owned submission ledger. Assign names during Gate 1 and update status at every design review.

| Artifact | Primary owner | Required cross-review | Gate | Submission role |
| --- | --- | --- | --- | --- |
| Team charter and final BRD | BI lead | Entire team | Gate 1 | Decision and scope contract |
| Lab 01 Store Health package | Report and dashboard designer | Evidence and source auditor | Gate 2 | Operations view, customer-voice view, and reality gap |
| Lab 02 Store Experience package | Evidence and source auditor | Report and dashboard designer | Gate 2 | Rating shape, coverage, and transferable-experience finding |
| Lab 03 Evidence Boundary package | Evidence and source auditor | BI lead | Gate 3 | Join audit, research queue, and store-month specification |
| Lab 04 Offer Pilot package | AI and governance reviewer | Evidence and source auditor | Gate 3 | Receipt funnel, candidate cells, and holdout design |
| BI report and evidence audit | Report and dashboard designer | BI lead and evidence auditor | Gate 4 | Integrated analytical experience and traceable claims |
| Capstone Journal | AI and governance reviewer | Every artifact owner | Gate 4 | Notes, task evidence, Copilot prompts, provenance, verification, contribution, and decision rights |
| Executive brief and presentation | BI lead | Entire team | Gate 5 | Final recommendation and executive defense |

The manifest describes ownership, not silos. Every owner must obtain the named cross-review before moving the artifact to the next gate.

##### What You Carry Forward

| Stage | Artifact earned | Used later for |
| --- | --- | --- |
| BRD | One-page team framing and accepted decision contract | Scope control and final acceptance check |
| Lab 01 | Simulated operations view, observed customer-voice view, and missing join | Reality-gap section |
| Lab 02 | Rating-shape comparison, coverage check, and bounded theme finding | Transferable-experience recommendation |
| Lab 03 | Grain audit, independent state views, research queue, and store-month specification | Evidence boundary and data investment |
| Lab 04 | Receipt funnel, offer comparison, candidate cells, and holdout design | Pilot recommendation and guardrails |
| Synthesis | One finding per lab with its evidence class and boundary | Executive evidence snapshot |
| Executive brief | Recommendation, trade-offs, roadmap, risks, and refused claims | Final decision request |
| Capstone Journal | Working notes, task log, Copilot prompt log, provenance, verification, contributions, and human decision owners | Responsible AI and academic-integrity evidence |

##### Case Data Package

For the fictional engagement, the BI team receives six public-source or simulated extracts assigned to case roles:

| Case role | Data used in | What it supports |
| --- | --- | --- |
| Operations Analytics | Lab 01 | Store-health dashboard workflow |
| Customer Experience | Labs 01-03 | Customer voice, ratings, and feedback coverage |
| Store Development | Lab 03 | Store-network geography |
| People Analytics | Lab 03 | Workforce-event research prioritization |
| Marketing Analytics | Lab 04 | Campaign funnel and pilot design |

Use the fields and limits introduced by each lab. The package is assembled for independent educational use; do not alter source files or attempt to reconstruct missing operational history.

###### Evidence Labels

- **Observed:** Make bounded descriptive claims about the named sample, period, and selection mechanism.
- **Simulated:** Demonstrate a workflow or formulate a testable hypothesis. Never present it as actual Starbucks behavior.
- **Missing:** Name the absent field, grain, key, denominator, or time history needed for the next study.

###### Dataset Boundaries

- The extracts can support the lab questions, not a claim that customer feedback caused revenue decline or unionization.
- The extracts do not supply the store-month join required for a production Store Health system.

##### Phase 1: Understand the Business Request

###### Read

- [BRD](brd.md)

###### Outcome

Write a one-page team framing before opening Power BI:

- team charter: primary roles, artifact ownership, and cross-review responsibilities;
- decision owner: Maya Rodriguez;
- target users: store manager, district manager, and executive;
- decision to support: which analytics operating model should Starbucks fund;
- three competing options: Centralized, Self-Service, Governed Self-Service;
- one hypothesis: a governed store-health layer can improve local action without turning every team into an uncontrolled metric publisher;
- one disconfirming condition: the evidence shows no meaningful decision value beyond existing reporting or cannot be governed operationally.

###### Gate 1: Brief accepted

The team can explain its roles, the decision, users, evidence classes, and out-of-scope claims without opening a dataset.

##### Phase 2: Analyze the Evidence

Complete the labs in order. Each lab answers a different question and leaves an artifact for the final executive brief.

###### Lab 01: Store Health Simulation

[Open Lab 01](labs/01-store-health-dashboard.md)

**Question:** Which operational dimensions differ in the simulated ordering data, and what can an operations-only dashboard see?

**Use:** B3 simulated ordering data, then A2 observed NYC reviews as a reality check.

**Team artifact:** One simulated-driver visual, one observed customer-voice visual, and a one-sentence statement of the missing store-by-time join.

**Do not claim:** B3 represents real Starbucks operations or that its channel differences caused the A2 review change.

###### Lab 02: Same Brand, Different Experience

[Open Lab 02](labs/02-store-comparison.md)

**Question:** What differs between Broadway and the Reserve Roastery when rating shape and customer language are kept visible?

**Use:** A2 observed public reviews from two named NYC stores.

**Team artifact:** A distribution comparison with sample sizes, a language-treatment note, and one transferable experience element plus one non-transferable flagship cost.

**Do not claim:** The two stores establish a network effect or a parallel time trend.

###### Gate 2: Operations and customer signal accepted

The team can explain what the operations dashboard sees, what customer voice adds, why the sources cannot be joined, and which Lab 01-02 artifacts will enter the final brief.

###### Lab 03: Evidence Boundary

[Open Lab 03](labs/03-geographic-evidence-boundary.md)

**Question:** Where is public evidence dense enough to prioritize primary research, and where does the store join fail?

**Use:** A1 reviews, A3 2021 locations, and A6 election records.

**Team artifact:** Three independent state views, the 17-state classroom research queue, and a store-month data specification.

**Do not claim:** State co-presence proves same-store overlap, customer complaints caused unionization, or a composite risk score is valid.

###### Lab 04: Offer Completion

[Open Lab 04](labs/04-offer-completion.md)

**Question:** Which simulated offer and segment differences justify a randomized pilot rather than a production targeting rule?

**Use:** B1 simulated campaign data.

**Team artifact:** A corrected receipt-level funnel, an offer-type comparison, a segment-by-offer candidate view, and a holdout design.

**Do not claim:** Simulated demographic differences prove causation, fairness, or permission to exclude customers.

###### Gate 3: Evidence boundary and pilot accepted

The team has four lab artifacts and can state, for every headline number:

- source;
- grain;
- period;
- denominator;
- evidence class; and
- safe decision use.

##### Phase 3: Synthesize the Findings

Before writing the executive brief, build a synthesis table:

| Finding | Evidence | Decision implication | Boundary |
| --- | --- | --- | --- |
| B3 has visible channel differences | Simulated | Demonstrate local-driver workflow | Not actual Starbucks performance |
| A2 shows a Roastery decline | Observed two-store sample | Require customer voice in the design | Not a network trend or causal join |
| A1/A3/A6 share state coverage | Observed selected samples | Prioritize primary research | No verified store-level relationship |
| B1 changes the winning offer metric | Simulated | Test completion with a holdout | No production targeting rule |

Then answer the executive question:

> What should Starbucks fund now, what should store teams be allowed to do locally, and what evidence must be collected before stronger claims are made?

###### Required recommendation structure

1. Choose one operating model.
2. Name the core certified metrics.
3. Define what local teams may investigate or annotate.
4. Define Copilot guardrails and human review points.
5. Specify the missing store-month data model.
6. Name rollout risks and the first pilot boundary.

Before passing this gate, begin the [Capstone Journal](reports/journal.md). Record working notes, completed tasks, representative Copilot prompts and critiques, verification performed, and the human owner of each consequential decision.

###### Gate 4: Synthesis and design review accepted

The recommendation changes a decision, names a trade-off, survives a ship/iterate/discard review, and does not depend on a causal claim the evidence cannot support.

##### Phase 4: Create and Present the Executive Brief

[Open the Executive Brief Template](reports/executive-brief.md)

Use the blank report structure for the final team output. Fill it with the team's own visuals, recommendation, trade-offs, and refused claims.

###### Required report flow

1. **Decision first:** what Maya should approve.
2. **Evidence snapshot:** one finding per evidence class.
3. **Reality gap:** what the simulated operations layer can see versus what observed customer voice reveals.
4. **Evidence boundary:** where public data stops and primary research begins.
5. **Operating model:** certified metrics, local investigation, Copilot guardrails.
6. **Roadmap:** govern, instrument, pilot, scale.
7. **Refused claims:** what the team will not infer from the data.
8. **Decision request:** the specific approval needed from the executive.

###### Final presentation

Unless the instructor specifies another format, present a **7-minute executive briefing** followed by questions. Include:

- one recommendation;
- two trade-offs;
- three evidence-backed findings;
- one visual reality gap;
- one missing-data design;
- one refused claim; and
- one next action with an owner.

###### Gate 5: Executive defense accepted

The executive can answer these questions after the presentation:

- What should I approve?
- Why does the evidence support it?
- What is simulated or selected?
- What could still be wrong?
- Who acts next, and what will we measure?

##### Working Rules

- Keep observed, simulated, and missing evidence visibly separate.
- Do not replace pinned data for convenience.
- Do not join sources merely because they share a state or text field.
- Use rates with denominators, not raw counts alone.
- Treat Copilot as an analysis assistant, not an evidence source.
- Disclose material AI assistance and verify every AI-assisted calculation or claim.
- Preserve the strongest uncertainty that matters to the decision.
- Use the [BRD acceptance criteria](brd.md#8-acceptance-criteria) before presenting.


### BRD: Starbucks Store Health

#### BRD: Starbucks Store Health

**BRD means Business Requirements Document.** It is the request the student BI team receives before beginning analysis.

**Role-play boundary:** Maya Rodriguez and this engagement are fictional. The case is independently authored and is not affiliated with or endorsed by Starbucks Corporation.

**Decision owner:** Maya Rodriguez, fictional VP of Store Operations Analytics.

**Decision:** Recommend an analytics operating model that helps store teams act on customer, operational, and employee signals without creating metric chaos.

**Project format:** Teams of 3-4 unless the instructor approves another format.

##### Team Charter

Complete this charter before opening a dataset. A three-person team may combine the final two roles, but no person may approve their own evidence or AI-use record without cross-review.

| Team member | Primary role | Artifacts owned | Cross-review responsibility |
| --- | --- | --- | --- |
| [Name] | BI lead | BRD, synthesis, executive brief | Operating-model coherence and final decision chain |
| [Name] | Evidence and source auditor | Source audit, evidence labels, join decisions | Headline measures, denominators, and refused claims |
| [Name] | Report and dashboard designer | Model, measures, visuals, interaction, accessibility | Question-to-form and executive hierarchy |
| [Name, if applicable] | AI and governance reviewer | AI-use record, CSAR evidence, guardrails | Generated claims, fairness risks, and human decision boundaries |

**Decision owner:** Maya Rodriguez remains the fictional executive owner. The BI lead coordinates the team but does not replace the decision owner.

##### Business Context

Starbucks leaders receive customer, operations, marketing, and people signals through separate teams. The requested solution is a focused Store Health experience that helps a manager answer:

- What changed?
- Where is the change concentrated?
- What evidence supports the finding?
- What action is within my authority?
- What is still unknown?

The team must assess whether **Governed Self-Service** is a stronger operating model than centralized reporting or unrestricted self-service.

##### Users and Decisions

| User | Decision supported | Required view |
| --- | --- | --- |
| Store manager | What local condition needs attention today? | Store detail and operating drivers |
| District manager | Which locations need investigation or support first? | Comparable locations or research priorities |
| Executive | Which operating model and data investment should be funded? | Evidence summary, risks, and recommendation |

##### Case Data Package

For this fictional engagement, the course facilitator has assembled the following public-source and simulated extracts:

| Data package | Providing team | Used in | Essential boundary |
| --- | --- | --- | --- |
| Store operations sample | Operations Analytics | Lab 01 | Simulated; not actual store performance |
| Customer voice sample | Customer Experience | Labs 01-02 | Two named NYC stores; no operations join |
| Customer feedback geography | Customer Experience | Lab 03 | Location is not a verified store identifier |
| Store network snapshot | Store Development | Lab 03 | Historical snapshot; no shared key with other extracts |
| Workforce event extract | People Analytics | Lab 03 | State-level planning only; not a causal employee-condition measure |
| Rewards campaign sample | Marketing Analytics | Lab 04 | Simulated; no production targeting rule |

Use the extracts as issued. Do not replace, merge, or normalize source files while completing the labs.

##### Required Analysis

1. Identify the operational dimensions visible in the store-health sample.
2. Compare those dimensions with customer voice and name the missing join.
3. Compare Broadway and Reserve Roastery rating shape, themes, and coverage limits.
4. Create a research queue without implying a false store-level relationship.
5. Identify campaign patterns that earn a randomized pilot, not direct deployment.
6. Specify the data and ownership required for a real Store Health system.

##### Required Report Package

###### Executive Evidence Summary

Show the decision, recommendation, one headline finding per evidence class, and the principal trade-off.

###### Store Health View

Show simulated operational drivers, satisfaction distribution, low-rating share, fulfillment context, and the limits of the operational extract.

###### Customer Voice Reality Check

Show the two-store rating comparison, yearly sample sizes, language-treatment limits, and the missing operational fields.

###### Evidence Boundary and Research Queue

Show independent state-level coverage, source grain, research priorities, and the missing store-month model.

###### Campaign Pilot

Show offer completion, valid segment comparisons, pilot candidates, and holdout guardrails.

###### AI Use and Provenance

Provide a completed Capstone Journal containing working notes, task-completion evidence, representative Copilot prompts and critiques, team contributions, data provenance, verification results, ship/iterate/discard decisions, and human approval boundaries.

##### Evidence Rules

- Label every finding as observed, simulated, or missing evidence.
- Use rates and their denominators, not raw counts alone.
- Preserve each source's grain and period.
- Do not create a store-level relationship from a shared state field.
- Do not present simulated findings as actual Starbucks behavior.
- Do not use Copilot output as evidence without checking the data and definitions.
- State what the evidence earns and what it does not earn.

##### Deliverables

1. A team charter and contribution statement naming primary ownership and cross-review responsibility.
2. A final BRD recording any accepted changes to the decision contract.
3. A Power BI report or equivalent interactive artifact.
4. A source-and-join audit describing grain, periods, denominators, selection mechanisms, and missing evidence.
5. A reality-gap comparison and store-month data specification.
6. A completed Capstone Journal with working notes, task evidence, representative Copilot prompts, provenance, and verification evidence.
7. A one-page executive recommendation comparing Centralized, Self-Service, and Governed Self-Service.
8. A 7-minute executive briefing with one recommendation, two trade-offs, one refused claim, and one next action with an owner.

##### Milestone Gates

| Gate | Evidence required | Exit condition |
| --- | --- | --- |
| **Gate 1: Team charter and BRD** | Roles, ownership, users, decision, question, comparison, success condition, and scope boundaries | The team can explain the decision without opening a dataset |
| **Gate 2: Operations and customer signal** | Labs 01-02, evidence labels, sample sizes, coverage limits, and reality gap | The team can state what the dashboard sees and what it cannot explain |
| **Gate 3: Evidence boundary and pilot** | Labs 03-04, grain audit, research queue, receipt funnel, and holdout design | The team refuses false joins and converts patterns into a fair test |
| **Gate 4: Synthesis and design review** | One finding per lab, operating-model choice, dashboard draft, AI-use record, and provenance statement | Every headline claim traces to a source and survives adversarial review |
| **Gate 5: Executive defense** | Complete submission package and rehearsal against the rubric | The audience can identify the approval, trade-off, risk, next owner, and revisit evidence |

##### Acceptance Criteria

The work is ready for presentation when:

- every required question has a visible answer or an explicit evidence gap;
- every page names its evidence class;
- every headline metric shows a source, period, and denominator;
- no visual implies a verified store-level relationship where none exists;
- no visual presents simulated data as actual Starbucks behavior;
- team ownership and cross-review are documented;
- material AI assistance, verification, and human decision rights are disclosed;
- the recommendation names governance, local flexibility, Copilot guardrails, ownership, and rollout risk; and
- the executive can identify the next decision, the owner, and the evidence needed to revisit it.

##### Out of Scope

- Proving that customer complaints caused revenue decline or unionization.
- Reconstructing private production history from the course extracts.
- Deploying a production workspace or targeting customers.
- Treating a synthetic correlation or Key Influencers result as causal evidence.

##### Starting Path

1. Read the Capstone Case, form a team, and assign primary and cross-review roles.
2. Download the required extracts from the [Dataset Package](dataset-package.md).
3. Complete Labs 01-04 in order.
4. Synthesize the findings in the Executive Brief.
5. Present the recommendation, trade-offs, and refused claims to the case evaluator.

*Data acknowledgement: the course package includes reviewed public-source material and explicitly simulated datasets. Technical lineage, licenses, redistribution status, and known limitations are maintained in the data-governance record.*


### Dataset Package

#### Dataset Package

Download the files for each lab before opening Power BI. This independent educational package combines public-source material with explicitly simulated extracts; it is not a Starbucks-provided data package.

> **Publication boundary:** The unlisted student and professor routes cache every hash-pinned teaching input for reproducibility. Anyone with a direct route URL can download the files; neither route appears in the main site navigation.

##### Capstone Data Contract

| ID | File or files | Lab | Evidence class | Allowed use | Do not use for |
| --- | --- | --- | --- | --- | --- |
| A1 | [reviews_data.csv](data/A1/reviews_data.csv) | Lab 03 | Observed selected sample | Reviewer-geography coverage and low-rating counts | Store location, network sentiment, or causal labor claims |
| A2 | [starbucks_ny_broadway.csv](data/A2/starbucks_ny_broadway.csv); [starbucks_ny_reserve_roastery.csv](data/A2/starbucks_ny_reserve_roastery.csv) | Labs 01-02 | Observed two-store sample | Rating shape, review coverage, and bounded theme comparison | Network trend or causal operations explanation |
| A3 | [startbucks.csv](data/A3/startbucks.csv) | Lab 03 | Observed 2021 snapshot | State-level store exposure | Current network census or deterministic join to A1/A6 |
| A6 | [sbwu_elections.csv](data/A6/sbwu_elections.csv) | Lab 03 | Observed selected workforce-event records | Election coverage and primary-research planning | Employee-condition measure or customer-cause inference |
| B1 | [portfolio.csv](data/B1/portfolio.csv); [profile.csv](data/B1/profile.csv); [transcript.csv](data/B1/transcript.csv) | Lab 04 | **SIMULATED** campaign | Receipt funnel, offer comparison, and pilot design | Actual targeting rule or demographic causation |
| B3 | [starbucks_customer_ordering_patterns.csv](data/B3/starbucks_customer_ordering_patterns.csv) | Lab 01 | **SIMULATED** operations | Dashboard workflow and grouped comparisons | Actual Starbucks performance or explanation of A2 reviews |

The upstream filename [startbucks.csv](data/A3/startbucks.csv) is intentional and hash-pinned. Do not rename or correct it in the source package. [manifest.json](data/manifest.json) is the authority for file names, schemas, row counts, licenses, and SHA-256 hashes.

**Team checkpoint:** The evidence and source auditor verifies hashes before analysis. Each artifact owner records source ID, grain, period, denominator, transformations, and human verifier in the Capstone Journal.

##### Lab 01: Store Health

Download all three files:

- [starbucks_customer_ordering_patterns.csv](data/B3/starbucks_customer_ordering_patterns.csv)
- [starbucks_ny_broadway.csv](data/A2/starbucks_ny_broadway.csv)
- [starbucks_ny_reserve_roastery.csv](data/A2/starbucks_ny_reserve_roastery.csv)

Use the operations sample for the dashboard workflow. Use the two customer-voice files for the reality check. Keep those layers separate: they do not form a causal join.

##### Lab 02: Store Comparison

Download both customer-voice files:

- [starbucks_ny_broadway.csv](data/A2/starbucks_ny_broadway.csv)
- [starbucks_ny_reserve_roastery.csv](data/A2/starbucks_ny_reserve_roastery.csv)

Add a `store` column before appending the files in Power Query.

##### Lab 03: Evidence Boundary

Download all three files:

- [reviews_data.csv](data/A1/reviews_data.csv)
- [startbucks.csv](data/A3/startbucks.csv)
- [sbwu_elections.csv](data/A6/sbwu_elections.csv)

Keep the extracts as separate tables. They support state-level planning, not a verified store-level relationship.

##### Lab 04: Offer Completion

Download all three files:

- [portfolio.csv](data/B1/portfolio.csv)
- [profile.csv](data/B1/profile.csv)
- [transcript.csv](data/B1/transcript.csv)

Use the provided files as issued. The lab explains the preparation required before calculating funnel and segment metrics.

##### Download Checklist

- Create one folder per lab on your machine.
- Keep each original file unchanged.
- Import the CSVs into Power BI from the lab folder.
- Follow the preparation steps in the corresponding lab document.
- Do not combine files across labs unless a lab explicitly asks you to.

##### Data Boundaries

The package contains observed and simulated extracts. Each lab identifies which is which and what conclusions are allowed. The files support the analysis request; they do not provide the full store-month history required for a production Store Health system.

##### Appendix: Data Credits

The course package includes reviewed public-source material from [Kaggle](https://www.kaggle.com/), curated for this independent educational case.

| Extract | Contributor |
| --- | --- |
| Customer feedback geography | Harshal H (`harshalhonde`) |
| NYC customer voice | Muqaddas Ejaz (`muqaddasejaz`) |
| Store network snapshot | KUKUROO3 (`kukuroo3`) |
| Workforce event extract | Brandon Conrady (`brandonconrady`) |
| Rewards campaign sample | Ihor Muliar (`ihormuliar`) |
| Store operations sample | Likitha Gedipudi (`likithagedipudi`) |

Kaggle and the contributors made the source material available for analysis. The course package provides the reviewed extracts used in each lab.


### Lab 01: Store Health

#### Lab 01: Store Health

**Business anchor:** [BRD](../brd.md)

**Question:** Can an operations dashboard explain the customer signal leaders need to act on?

**Decision:** Should the Store Health design require a customer-voice layer?

Start by treating the operations dashboard as a useful but incomplete management tool. Your job is not to prove why customer ratings changed. Your job is to show what the operations data can prioritize, what customer voice adds, and what measurement is missing between them.

##### Capstone Position

**Primary owner:** Report and dashboard designer. \
**Cross-review:** Evidence and source auditor. \
**Gate:** Gate 2, Operations and customer signal. \
**Evidence class:** SIMULATED operations + OBSERVED two-store customer voice + MISSING store-by-time connection. \
**Capstone artifact:** Store Health page, customer-voice page, reality-gap slide, and missing-measurement recommendation.

This lab opens the capstone evidence chain. It establishes why an operations-only dashboard is useful but insufficient.

> **Answer-key boundary:** Do not consult Worked Solutions while producing this artifact unless the instructor explicitly authorizes a comparison exercise.

##### Use These Files

| Layer | Files | What it represents |
| --- | --- | --- |
| Operations | [starbucks_customer_ordering_patterns.csv](../data/B3/starbucks_customer_ordering_patterns.csv) | Simulated orders, channels, fulfillment, and satisfaction |
| Customer voice | [starbucks_ny_broadway.csv](../data/A2/starbucks_ny_broadway.csv) and [starbucks_ny_reserve_roastery.csv](../data/A2/starbucks_ny_reserve_roastery.csv) | Observed reviews from two NYC stores |

The two layers do **not** join. Do not merge them or use one to explain the other.

##### What You Need to Produce

1. A one-page Store Health dashboard using the operations extract.
2. A customer-voice comparison for Broadway and Reserve Roastery.
3. One reality-gap slide: what the dashboard can see, what customer voice adds, and the missing data needed to connect them.
4. A one-sentence recommendation: should customer voice be required in the Store Health design?

##### Four Steps

###### 1. Build the Operations View

Import the store operations file into your BI tool. In Power BI, choose **Get Data > Text/CSV**. Create:

- a large-number card for average satisfaction (often called a KPI);
- a satisfaction distribution;
- satisfaction by `order_channel`;
- satisfaction by `order_ahead`.

Use the optional **Key Influencers** visual in Power BI, or compare the averages in the channel and order-ahead charts. You are looking for the biggest difference that is visible in the data, not a proof of cause.

Keep the page simple: one headline KPI, one distribution, and two driver comparisons. A manager should be able to identify the strongest visible difference in a few seconds.

**Expected signal:** Mobile App satisfaction is about 3.86; Drive-Thru is about 3.44. Treat this as a simulated workflow finding, not actual Starbucks performance.

###### 2. Build the Customer-Voice View

Import the Broadway and Reserve Roastery files. In each file, add a new column named `store`: enter `Broadway` for every Broadway row and `Roastery` for every Roastery row. Then append the two tables, meaning stack one table's rows below the other. Convert `publishedDate` to a date field so you can group reviews by year.

Create:

- rating distribution by store;
- mean rating by store with sample size;
- Reserve Roastery mean rating by year, with yearly review count.

Use a separate visual for the yearly Roastery trend. Do not combine the two stores into one trend line because their review histories do not cover the same periods.

**Expected signal:** Reserve Roastery moves from 4.35 in 2019 to 3.20 in 2023. Show the number of reviews beside every yearly value. That count is the denominator: it tells the executive how much evidence supports the average.

###### 3. Compare the Layers Without Joining Them

Put these two numbers on one slide:

- largest simulated operations gap: about 0.41 points;
- observed Roastery decline, 2019 to 2023: about 1.15 points.

The point of this slide is contrast, not calculation. Place the two numbers side by side, label one **SIMULATED OPERATIONS** and the other **OBSERVED CUSTOMER VOICE**, and add the missing store-by-time connection underneath.

Use this caption or a tighter version of it:

> The operations view ranks the variables it contains, but the observed customer signal is larger and cannot be explained without a store-by-time connection to customer voice.

###### 4. Recommend the Missing Measurement

Specify the minimum future table Starbucks needs:

- stable store ID;
- time period;
- transactions or visits;
- complaint volume and theme;
- staffing and operating conditions.

Write the recommendation as an investment request: customer voice should become a required layer, but only a stable store-by-time table can test which operating conditions explain it. A store-by-time table has one row for the same store during the same period, allowing the two departments' measures to be compared fairly.

Conclude whether customer voice should be a required Store Health layer.

##### Evidence Rules

- Label the operations view **SIMULATED**.
- Label the reviews **OBSERVED CUSTOMER VOICE: TWO NYC STORES**.
- Do not claim that order channel, fulfillment time, or any simulated variable caused the review decline.
- Do not present the Roastery trend as a network-wide Starbucks trend.
- Use sample size with every comparison that depends on reviews.

##### Deliverable Checklist

- [ ] Operations dashboard contains satisfaction, channel, and order-ahead views.
- [ ] Customer-voice view identifies both stores and shows review counts.
- [ ] Reality-gap slide compares 0.41 and 1.15 without causal language.
- [ ] Recommendation names the missing store-by-time measurement.
- [ ] Every visual is labeled observed, simulated, or missing evidence.

##### Journal Checkpoint

- In the [Capstone Journal](../reports/journal.md), record working notes, completed steps, source IDs B3 and A2, transformations, measures, filters, periods, denominators, and any Copilot prompts used.
- If AI assists with a formula, visual choice, or narrative, preserve the representative prompt, first output, critique, refinement, and final disposition.
- The evidence and source auditor independently verifies the 0.41 and 1.15 comparisons and confirms that no causal join was created.
- Apply a ship/iterate/discard decision to any AI-generated explanation of the channel gap or review decline.

##### Gate Contribution

Lab 01 does not pass Gate 2 alone. It contributes the operations view, observed customer signal, and missing-join statement. Lab 02 must add rating shape, coverage, and bounded theme evidence before the team requests Gate 2 review.

##### Handoff to the Executive Brief

Carry forward:

- the strongest operations finding;
- the Roastery decline with its sample years;
- the missing store-by-time join; and
- your recommendation to require customer voice in the design.

Proceed to [Lab 02](02-store-comparison.md) after completing this checklist.


### Lab 02: Same Brand, Different Experience

#### Lab 02: Same Brand, Different Experience

**Business anchor:** [BRD](../brd.md)

**Question:** What does the Reserve Roastery do differently from Broadway in the customer's voice?

**Decision:** Which experience elements are worth testing in conventional stores?

This lab is about comparing experience, not declaring a winner. The Roastery is a flagship format and Broadway is a conventional location. Look for what customers reward in each setting, then separate transferable service ideas from expensive or polarizing features.

##### Capstone Position

**Primary owner:** Evidence and source auditor. \
**Cross-review:** Report and dashboard designer. \
**Gate:** Gate 2, Operations and customer signal. \
**Evidence class:** OBSERVED customer voice from two named NYC stores. \
**Capstone artifact:** Rating-shape comparison, review-coverage check, language-treatment record, and transferable-experience recommendation.

This lab deepens the signal introduced in Lab 01. It shows why an average can hide polarization, sparse periods, and differences in what customers value.

> **Answer-key boundary:** Do not consult Worked Solutions while producing this artifact unless the instructor explicitly authorizes a comparison exercise.

##### Use These Files

- [starbucks_ny_broadway.csv](../data/A2/starbucks_ny_broadway.csv)
- [starbucks_ny_reserve_roastery.csv](../data/A2/starbucks_ny_reserve_roastery.csv)

Add a `store` column before appending the files. Treat the result as two named NYC stores, not a network sample.

##### What You Need to Produce

1. A rating-distribution comparison with a sample size for each store.
2. A review-volume-by-year chart for each store.
3. A short theme comparison using one consistent language treatment.
4. One transferable experience element and one feature that should not be scaled.

##### Four Steps

###### 1. Build One Review Table

In each file, add a new `store` column: use `Broadway` in the Broadway file and `Roastery` in the Roastery file. Append the files, meaning stack their rows into one review table. Convert `publishedDate` into a date and `rating` into a whole number.

Before building visuals, confirm that both store labels, dates, and ratings loaded correctly. Keep the original review text because you will use it for the theme comparison.

###### 2. Compare Rating Shape

Show ratings as a percentage of each store's reviews, not raw counts.

Use a 100% stacked bar or grouped percentage chart. A 100% stacked bar makes each store equal width, so a small store's ratings are not visually hidden by a larger review count. The question is whether the mix of ratings differs, not which store collected more reviews.

**Expected signal:** The mean gap is small, 3.92 versus 4.09. The stronger result is the shape: Roastery has about 56% five-star reviews versus 38% for Broadway, but also a slightly higher one-star share.

###### 3. Check Time Coverage

Show review count by year for each store before building a trend chart.

Treat this as a quality check. Use $N=10$ as the minimum annual review count for a reliable trend point. Keep thinner years visible in the coverage chart and show their means as hollow warning points. Use dashed connectors only between adjacent observed years when at least one endpoint is thin, and label those segments directional. Reserve solid lines for adequate samples and break every line across missing years rather than smoothing, interpolating, or compressing the gap.

**Expected signal:** Broadway has $N=5$ in 2022 and 2023 and $N=1$ in 2024; Roastery has $N=4$ in its partial opening year, 2018, and $N=1$ in 2024. The representative Roastery window, 2019-2023, has $N=13$ to $156$ per year. Do not create a parallel Broadway-versus-Roastery time trend.

###### 4. Compare Themes Carefully

Use the same language treatment for both stores. Compare themes in five-star and one-star reviews.

Choose and declare one language treatment before reading the results:

- analyze English-like reviews only;
- translate both stores first; or
- apply the provided corpus-informed multilingual keyword lexicon to normalized title and text, keep all reviews in the denominator, and disclose that it does not detect language, translate text, infer sentiment, or provide complete multilingual coverage.

State the rule, included review count, and limitation on the visual so the executive knows what portion of the customer voice is represented.

Treat the following as hypotheses to test through rating-stratified or human-coded review, not as findings from mention counts alone:

- whether Roastery ambiance, space, and experience mentions are positive;
- whether Broadway staff and welcome mentions are positive; and
- whether wait and staff-interaction mentions describe shared failures.

##### Evidence Rules

- Use within-store percentages for rating comparisons.
- Do not claim the Roastery is universally better because its mean rating is higher.
- Do not use the two stores to establish a network effect or parallel trend.
- Label any theme analysis with its language-treatment limitation.
- Keep each store's review count and period visible.

##### Deliverable Checklist

- [ ] Rating-shape visual uses within-store percentages.
- [ ] Review coverage by year explains why a parallel trend is invalid.
- [ ] Theme comparison uses one language rule for both stores.
- [ ] Recommendation names one validated behavior to test, or requests validation before naming one, plus one scalability limit.
- [ ] Every conclusion is scoped to the two observed stores.

##### Journal Checkpoint

- In the [Capstone Journal](../reports/journal.md), record working notes, completed steps, the appended-table transformation, review periods, sample sizes, language treatment, theme rules, and any Copilot prompts used.
- If AI assists with theme extraction or translation, disclose the tool and model, preserve one representative CSAR cycle, and test the output against manually reviewed examples from both stores.
- The report and dashboard designer verifies that within-store percentages, time coverage, labels, and visual encodings match the question.
- Discard any generated claim that turns two stores into a network conclusion or treats a keyword scan as complete sentiment analysis.

##### Gate Contribution

Lab 02 completes the evidence required for Gate 2. Before requesting review, combine its rating-shape and coverage findings with Lab 01's operations view, reality gap, and missing store-by-time join.

##### Handoff to Lab 03

Carry forward:

- the difference between a mean and a distribution;
- the time-coverage limitation; and
- the rule that a common brand does not create a comparable time series.

Proceed to [Lab 03](03-geographic-evidence-boundary.md) after completing this checklist.


### Lab 03: Evidence Boundary

#### Lab 03: Evidence Boundary

**Business anchor:** [BRD](../brd.md)

**Question:** Where is the evidence strong enough to justify primary research?

**Decision:** Which states should receive follow-up, and what data must Starbucks collect before testing a labor-and-customer-experience relationship?

This lab teaches a critical BI discipline: a shared column does not guarantee a valid relationship. The sources all mention geography, which makes a direct join tempting. Your task is to make the useful state-level planning view while refusing the false store-level story.

##### Capstone Position

**Primary owner:** Evidence and source auditor. \
**Cross-review:** BI lead. \
**Gate:** Gate 3, Evidence boundary and pilot. \
**Evidence class:** OBSERVED selected samples with MISSING stable store-level relationship. \
**Capstone artifact:** Grain-and-join audit, three independent state views, 17-state research queue, and store-month data specification.

This lab tests whether the team can stop an authoritative-looking analysis when the relationship is computable but not interpretable.

> **Answer-key boundary:** Do not consult Worked Solutions while producing this artifact unless the instructor explicitly authorizes a comparison exercise.

##### Use These Files

- [reviews_data.csv](../data/A1/reviews_data.csv)
- [startbucks.csv](../data/A3/startbucks.csv)
- [sbwu_elections.csv](../data/A6/sbwu_elections.csv)

Keep the three files separate. They share state-level geography, not a verified store key.

##### What You Need to Produce

1. A source-and-join audit.
2. Three aligned state views, one per source.
3. A 17-state research queue.
4. A one-row-per-store-per-month data specification for the next study.

##### Four Steps

###### 1. Audit the Grain

For each source, complete this table before creating a visual: **one row represents**, **location means**, **time period**, and **verified store ID available?**. In data work, this row definition is called the *grain*.

Write this audit before creating any chart. It is the control that prevents Power BI from producing an authoritative-looking but invalid result.

**Expected result:**

- A1 location belongs to the reviewer record, not a verified store;
- A3 identifies stores through `storeNumber`;
- A6 identifies filed elections without A3's stable store key.

Record the join verdict: **state-level coverage only**.

###### 2. Build Independent State Aggregates

Create one state-level table per source. An aggregate is simply a summary table with one row per state:

- A1: review count, rated review count, low-rating count;
- A3: store count;
- A6: election count and union-win count.

Use the same state code and sort order in every output. The views should be easy to compare, but they should remain separate measures.

Do not divide A1 reviews by A3 stores and call the result a complaint rate.

###### 3. Create the Research Queue

Create a simple list of state codes and place each source's counts beside the matching state. This is a side-by-side planning table, not a merge of individual reviews, stores, and elections.

Show the research-ready rule in the report exactly as written. The queue selects where to investigate first; it does not rank states by employee or customer risk.

Use the classroom rule:

```text
ResearchReady = RatedReviews >= 10 AND Elections >= 5
```

**Expected result:** 17 states clear the rule. This is a workload rule for follow-up, not a risk score or significance test.

###### 4. Specify the Missing Study

Define a store-month table with:

- stable store ID and period;
- transactions or visits;
- complaint volume and theme;
- staffing, turnover, and employee conditions;
- organizing milestones and market controls.

This is the action step. Translate the failed join into a concrete data request that an operations, customer experience, and people analytics team could actually build. “Store-month” means one row for one store during one month.

##### Evidence Rules

- Label A1 geography as **reviewer-reported**, not store location.
- Keep counts, periods, and denominators visible.
- Show separate state views, not a composite labor/customer risk score.
- Do not fuzzy-match free-text location into a causal dataset.
- State co-presence is not same-store overlap.

##### Deliverable Checklist

- [ ] Source-and-join audit identifies grain, key, period, and selection mechanism.
- [ ] Three state views remain separate.
- [ ] Research queue uses the stated classroom rule and shows both denominators.
- [ ] Final recommendation requests primary research, not a causal conclusion.
- [ ] Store-month specification names stable store ID and time.

##### Journal Checkpoint

- In the [Capstone Journal](../reports/journal.md), record working notes, completed steps, source IDs A1/A3/A6, grain, period, geography meaning, selection mechanism, aggregate logic, join verdict, and any Copilot prompts used.
- If AI proposes fuzzy matching, a composite score, or a causal story, preserve that output and document why the team iterated or discarded it.
- The BI lead verifies that the research queue uses the stated classroom rule and that each denominator remains visible.
- Name the human owner who approves the primary-research recommendation and the evidence required before any store-level test.

##### Gate Contribution

Lab 03 supplies the evidence-boundary half of Gate 3. Lab 04 must add a bounded pilot with valid analysis population, thresholds, and holdouts before the team requests Gate 3 review.

##### Handoff to Lab 04

Carry forward the rule that a computable relationship is not automatically an interpretable relationship.

Proceed to [Lab 04](04-offer-completion.md) after completing this checklist.


### Lab 04: Offer Completion

#### Lab 04: Offer Completion

**Business anchor:** [BRD](../brd.md)

**Question:** Which offer and segment patterns justify a randomized pilot?

**Decision:** Should Marketing lead with discount or BOGO, and what should be tested before any targeting rule is deployed?

This lab separates attention from action. A campaign can generate many views without generating the behavior the business values. Build the funnel first, then use the segment analysis only to design a fair, measurable pilot.

##### Capstone Position

**Primary owner:** AI and governance reviewer. \
**Cross-review:** Evidence and source auditor. \
**Gate:** Gate 3, Evidence boundary and pilot. \
**Evidence class:** SIMULATED campaign behavior and MISSING production experiment evidence. \
**Capstone artifact:** Receipt-level funnel, paid-offer comparison, pilot-candidate table, and randomized holdout design.

This lab closes the analytical sequence by converting a descriptive pattern into a test with an owner, comparison group, threshold, and exit condition.

> **Answer-key boundary:** Do not consult Worked Solutions while producing this artifact unless the instructor explicitly authorizes a comparison exercise.

##### Use These Files

- [portfolio.csv](../data/B1/portfolio.csv)
- [profile.csv](../data/B1/profile.csv)
- [transcript.csv](../data/B1/transcript.csv)

This is a simulated campaign extract. It supports a pilot design, not a production targeting decision.

##### What You Need to Produce

1. A receipt-level funnel: received, viewed, completed.
2. A view-rate and completion-rate comparison by offer type.
3. A segment-by-offer comparison using valid demographic records.
4. A randomized pilot design with a holdout in every tested segment.

##### Four Steps

###### 1. Prepare the Data

- Remove `Unnamed: 0` columns.
- Exclude `age = 118` from demographic analysis.
- Convert `transcript.value` from text into usable fields.
- Create one `offer_id` field by using `offer id` when present and `offer_id` otherwise.

Check the row count after each preparation step. The goal is not to remove inconvenient records; it is to create a trustworthy analysis population and preserve a clear explanation of every exclusion. If you use Power Query, split the `value` text into fields before joining it to the offer catalog.

###### 2. Build the Receipt-Level Funnel

Create one row per person and offer. Call this one row an **offer receipt**: one person receiving one specific offer. Mark each receipt with:

- `received_flag`;
- `viewed_flag`;
- `completed_flag`;
- `offer_type`.

Use `offer received` as the funnel denominator. A denominator is the total number of eligible receipts used to calculate a rate. A person cannot complete an offer they were never sent, so event counts alone are not enough.

Exclude informational offers when comparing completion; they cannot complete by design.

###### 3. Compare Offer Performance

Show view rate and completion rate by paid offer type.

Put both rates on the same visual or immediately adjacent visuals. The contrast between attention and completion is the central management insight in this lab.

**Expected signal:** BOGO has higher view rate, about 85% versus 72%. Discount has higher completion rate, about 61% versus 54%.

Use completion as the action metric. Views are attention, not the decision outcome.

###### 4. Design the Pilot

Join valid demographics to paid-offer receipts. Compare completion by income bucket and offer type.

Use right-inclusive income buckets and state the boundary convention: Under $40K means $x \le 40{,}000$; $40K-$60K means $40{,}000 < x \le 60{,}000$; continue the same pattern through $100K; and $100K+ means $x > 100{,}000$.

Only promote cells that meet both the difference and sample-size rule. The output is a pilot candidate list, not a list of customers to target or exclude immediately. A holdout is a randomly selected comparison group that keeps the current approach, so the pilot result can be measured fairly.

A cell becomes a pilot candidate only when:

```text
absolute difference from paid-offer baseline >= 10 points
AND receipt count >= 500
```

Keep a business-as-usual holdout in every tested segment.

##### Evidence Rules

- Label every campaign visual **SIMULATED**.
- Do not use `age = 118` as a real age.
- Do not use informational offers in a completion-rate comparison.
- Do not claim income, gender, or age causes completion.
- Do not turn a segment difference into an exclusion or production targeting rule.

##### Deliverable Checklist

- [ ] Funnel has one row per receipt and uses the coalesced offer ID.
- [ ] View and completion rates are shown side by side.
- [ ] Informational offers are excluded from the completion comparison.
- [ ] Segment cells show receipt count and difference from baseline.
- [ ] Pilot design includes a holdout in every segment.

##### Journal Checkpoint

- In the [Capstone Journal](../reports/journal.md), record working notes, completed steps, source B1, preparation row counts, offer-key logic, receipt grain, exclusions, baselines, thresholds, calculations, and any Copilot prompts used.
- If AI assists with parsing, segmentation, or pilot language, preserve one representative CSAR cycle and independently verify the output against the prepared receipt table.
- The evidence and source auditor verifies informational-offer exclusion, sample-size thresholds, differences from baseline, and the absence of causal demographic claims.
- State which targeting, exclusion, staffing, or automated decisions remain under human authority.

##### Gate Contribution

Lab 04 completes Gate 3 when combined with Lab 03. Before requesting review, the team must be able to defend the failed join, research queue, receipt population, pilot threshold, holdout, and prohibited inferences.

##### Handoff to the Executive Brief

Carry forward:

- completion as the primary paid-offer metric;
- the discount-versus-BOGO result;
- the pilot threshold and holdout design; and
- the rule that simulated segment differences do not justify production targeting.

Proceed to [Phase 3 synthesis](../README.md#phase-3-synthesize-the-findings), then complete the [Executive Brief](../reports/executive-brief.md) and [Capstone Journal](../reports/journal.md).


### Executive Brief Template

#### Executive Brief Template

Complete this brief after finishing Labs 01-04. Replace every bracketed prompt with the team's own analysis. Do not copy the Worked Solutions package into this submission.

Keep the executive brief to one page when exported. Submit the BI artifact, evidence audit, and completed [Capstone Journal](journal.md) as separate attachments.

| | |
| --- | --- |
| **Audience** | [Name the fictional decision owner and operating audience] |
| **Decision** | [State the approval or choice this brief requests] |
| **Recommendation** | [Choose Centralized, Self-Service, Governed Self-Service, or another defensible option] |
| **Evidence status** | [Summarize the observed, simulated, and missing evidence used] |
| **Source contract** | [Link the BRD and identify the lab outputs carried forward] |

##### Big Idea

> [Write one sentence that states what should change, why the evidence earns that change, and the principal boundary.]

##### Decision Requested

[State the specific approval, investment, pilot, or governance decision required. Name the owner and intended timing.]

##### Options And Trade-Offs

| Option | What it solves | What it risks | Evidence required |
| --- | --- | --- | --- |
| Centralized analytics | [Complete] | [Complete] | [Complete] |
| Unrestricted self-service | [Complete] | [Complete] | [Complete] |
| Governed self-service | [Complete] | [Complete] | [Complete] |

##### Evidence Snapshot

Carry forward one decision-relevant result from each lab. Every row must identify its evidence class and safe use.

| Lab | Finding | Source, period, and denominator | Evidence class | What it earns | What it does not earn |
| --- | --- | --- | --- | --- | --- |
| Lab 01: Store Health | [Complete] | [Complete] | [Observed, simulated, or missing] | [Complete] | [Complete] |
| Lab 02: Store Comparison | [Complete] | [Complete] | [Observed, simulated, or missing] | [Complete] | [Complete] |
| Lab 03: Evidence Boundary | [Complete] | [Complete] | [Observed, simulated, or missing] | [Complete] | [Complete] |
| Lab 04: Offer Completion | [Complete] | [Complete] | [Observed, simulated, or missing] | [Complete] | [Complete] |

##### Visual 1: Store Health

[Insert one visual showing what the simulated operations layer can prioritize. Label it **SIMULATED** and state its denominator.]

**Interpretation:** [Explain the visible comparison without causal language.]

##### Visual 2: Customer Voice Reality Check

[Insert one observed customer-voice visual. Name the stores, period, sample size, and selection limitation.]

**Interpretation:** [Explain what the observed sample adds and why it cannot be joined causally to the operations simulation.]

##### Visual 3: Evidence Boundary

[Insert the source-and-join audit or research queue. Keep each source's grain and denominator visible.]

**Interpretation:** [Explain where public evidence stops and why primary research is the next action.]

##### Proposed Operating Model

###### Certified evidence layer

- [Name the measures that require central definitions and owners.]
- [State how source, grain, period, denominator, and evidence class travel with each measure.]

###### Local investigation layer

- [State what store and district teams may filter, annotate, or investigate.]
- [State what local teams may not redefine or automate.]

###### Copilot guardrails

- [State where Copilot may summarize or suggest questions.]
- [Name the consequential decisions that require human review.]
- [Describe the audit trail required for generated answers and actions.]

##### Missing Store-Month Measurement

Specify one row per store per month.

| Required field | Business purpose | Owner | Quality check |
| --- | --- | --- | --- |
| Stable store ID and month | [Complete] | [Complete] | [Complete] |
| Transactions or visits | [Complete] | [Complete] | [Complete] |
| Complaint volume and theme | [Complete] | [Complete] | [Complete] |
| Staffing and operating conditions | [Complete] | [Complete] | [Complete] |
| Workforce conditions or milestones | [Complete] | [Complete] | [Complete] |
| Market and store-format controls | [Complete] | [Complete] | [Complete] |

##### Pilot Roadmap

| Phase | Outcome | Owner | Exit test |
| --- | --- | --- | --- |
| Govern | [Complete] | [Complete] | [Complete] |
| Instrument | [Complete] | [Complete] | [Complete] |
| Pilot | [Complete] | [Complete] | [Complete] |
| Scale or stop | [Complete] | [Complete] | [Complete] |

##### Risks And Mitigations

| Risk | Why it matters | Mitigation | Revisit signal |
| --- | --- | --- | --- |
| [Risk 1] | [Complete] | [Complete] | [Complete] |
| [Risk 2] | [Complete] | [Complete] | [Complete] |
| [Risk 3] | [Complete] | [Complete] | [Complete] |

##### Claims This Brief Refuses

- [Name one causal claim the evidence cannot support.]
- [Name one network-wide claim the selected samples cannot support.]
- [Name one simulated finding that cannot be presented as actual behavior.]
- [Name one targeting, workforce, or operational decision that needs additional evidence and human review.]

##### Final Executive Sentence

> [Write one sentence naming the decision, evidence, trade-off, and immediate next owner.]

##### Source Trail

- [BRD](../brd.md)
- [Lab 01: Store Health](../labs/01-store-health-dashboard.md)
- [Lab 02: Store Comparison](../labs/02-store-comparison.md)
- [Lab 03: Evidence Boundary](../labs/03-geographic-evidence-boundary.md)
- [Lab 04: Offer Completion](../labs/04-offer-completion.md)

##### Completion Gate

- [ ] The first screen states the decision and recommendation.
- [ ] Every headline number names its source, period, and denominator.
- [ ] Every finding is labeled observed, simulated, or missing evidence.
- [ ] The brief separates useful comparison from causal explanation.
- [ ] The proposed operating model names ownership and local authority.
- [ ] Copilot assistance has explicit human-review boundaries.
- [ ] The missing store-month model is concrete enough to build.
- [ ] Risks include observable revisit signals.
- [ ] Refused claims are explicit.
- [ ] The final sentence names the next owner and action.

##### Submission Attachments

- [ ] BI report or equivalent interactive artifact.
- [ ] Source-and-join audit and missing store-month specification.
- [ ] Team charter and contribution statement.
- [ ] [Capstone Journal](journal.md).
- [ ] Executive presentation and speaking plan.

##### Package Assembly Order

Submit the capstone as one named package in this order:

1. Executive brief.
2. BI report or equivalent artifact link.
3. Evidence audit and missing store-month specification.
4. Completed Capstone Journal.
5. Executive presentation.

Use consistent team, project, and artifact names across every file. The BI lead verifies package completeness; the evidence and source auditor verifies the source trail; the AI and governance reviewer verifies disclosure and human decision boundaries.

##### Handoff to Gate 5

The package is ready for executive defense when every attachment is present, every headline claim appears in the evidence audit, every material Copilot use appears in the Journal, and every team member can explain one reason the recommendation could be wrong.


### Capstone Journal

#### Capstone Journal

Use this journal throughout the capstone. Record working notes as decisions are made, list what the team completed and how, preserve the Copilot prompts that materially shaped the work, and document verification before submission.

Do not include credentials, private data, reviewer names, or raw review text in prompts, notes, or screenshots.

##### Project Record

| Field | Entry |
| --- | --- |
| Team name | [Complete] |
| Team members | [Complete] |
| Submission date | [Complete] |
| BI artifact | [File name or URL] |
| Executive brief | [File name or URL] |
| Copilot and other AI tools used | [Tool, model, and access surface] |

##### Working Notes

Add concise notes while you work. Capture uncertainty and next actions, not polished retrospective prose.

| Date | Lab or gate | Question, observation, or decision | Evidence consulted | Next action |
| --- | --- | --- | --- | --- |
| [Date] | [Lab / Gate] | [Note] | [Dataset, visual, source, or meeting] | [Next step] |
| [Date] | [Lab / Gate] | [Note] | [Complete] | [Complete] |

##### Task Completion Log

List what the team completed and how. Link each task to evidence another person can inspect.

| Task | Owner | What we completed and how | Evidence or file | Status | Human verifier |
| --- | --- | --- | --- | --- | --- |
| [Task] | [Name] | [Steps, method, and material choices] | [File, visual, query, or calculation] | [Not started / in progress / complete] | [Name] |
| [Task] | [Name] | [Complete] | [Complete] | [Status] | [Name] |

##### Copilot Prompt Log

Record representative, decision-relevant Copilot use rather than every autocomplete or spelling correction. Paste the exact Copilot prompt or instruction, then record what the team accepted, changed, or rejected.

| Date | Lab or task | CSAR stage | Exact Copilot prompt or instruction | Output summary | Human critique and revision | Disposition | Verifier |
| --- | --- | --- | --- | --- | --- | --- | --- |
| [Date] | [Lab / task] | [Crystallize / Scope / Assemble / Refine] | [Prompt without credentials or raw review text] | [What Copilot produced] | [What was wrong, missing, or changed] | [Used / modified / rejected] | [Name] |
| [Date] | [Complete] | [Stage] | [Prompt] | [Summary] | [Critique] | [Disposition] | [Name] |

At least one sequence of entries must show a complete CSAR cycle: the initial question, scoped evidence, first output, human critique, refined prompt, and final disposition.

##### Data and Claim Notes

Track provenance and analytical boundaries for every material artifact or claim.

| Artifact or claim | Source | Grain and period | Evidence class | Transformation or measure | Boundary or refused claim | Human verifier |
| --- | --- | --- | --- | --- | --- | --- |
| [Headline finding] | [Dataset ID and file] | [Complete] | [Observed / simulated / missing] | [Complete] | [What the evidence does not prove] | [Name] |
| [Visual or measure] | [Complete] | [Complete] | [Complete] | [Complete] | [Complete] | [Name] |

##### Verification Log

| Output checked | Verification performed | Result | Material revision | Human decision owner |
| --- | --- | --- | --- | --- |
| [Measure or formula] | [Recalculation, source check, or test] | [Pass / fail] | [Complete] | [Name] |
| [Visual] | [Question-to-form, scale, denominator, and accessibility review] | [Pass / fail] | [Complete] | [Name] |
| [Narrative claim] | [Source, denominator, and causality review] | [Pass / fail] | [Complete] | [Name] |

##### Ship, Iterate, or Discard Decisions

| Artifact or Copilot output | Decision | Reason | Required next action |
| --- | --- | --- | --- |
| [Artifact] | [Ship / iterate / discard] | [One-sentence rationale] | [Complete] |
| [Artifact] | [Decision] | [Reason] | [Next action] |

##### Team Contribution Statement

Every member must own work and cross-review another member's work.

| Team member | Primary role | Tasks and artifacts owned | Cross-review performed | Decision defended |
| --- | --- | --- | --- | --- |
| [Name] | [BI lead / evidence auditor / report designer / AI-governance reviewer] | [Complete] | [Complete] | [Complete] |
| [Name] | [Role] | [Complete] | [Complete] | [Complete] |
| [Name] | [Role] | [Complete] | [Complete] | [Complete] |
| [Name, if applicable] | [Role] | [Complete] | [Complete] | [Complete] |

##### Responsible AI and Decision Boundaries

- **Fairness risk reviewed:** [State whether segment, demographic, workforce, or access differences could create harm or exclusion.]
- **Human review required for:** [Name staffing, targeting, workforce, or automated decisions that AI may not authorize.]
- **Disclosure provided to audience:** [State how AI assistance and simulation are disclosed.]
- **Prohibited inference refused:** [Name at least one causal, network-wide, or individual-level claim the team rejected.]
- **Revisit evidence:** [State what new evidence would change the decision.]

##### Final Journal Check

Before executive defense:

- confirm completed tasks link to inspectable evidence;
- confirm representative Copilot prompts and human critiques are recorded;
- confirm every headline claim has a source, boundary, and human verifier;
- confirm every material AI-assisted output appears in the verification log;
- confirm each team member's contribution and cross-review are complete;
- confirm no prompt, note, or screenshot exposes credentials, private data, or raw review text; and
- link the completed journal from the final submission package.

> We verified the calculations and claims used in the capstone. Copilot output was treated as a draft or analytical aid, not as evidence. The named human owners approved the final recommendation and its decision boundaries.
