Green Run, Red Data
Orchestration status, SLA performance and data-quality evidence disagree at different grains. Build the release gate before closing the incident.
- ROLE
- Analytics Reliability Engineer
- TIMEBOX
- 75–105 minutes
- COMPLEXITY
- Intermediate / 3 of 5
- ROTATION
- Brief 14 of 16
The decision has already reached your desk.
The platform scorecard is largely green on orchestration status, and Operations wants to close the data incident. Downstream analysts report that some successful runs still produced questionable assets.
Pipeline executions, retries and data-quality checks have different grains. Successful runs can miss their SLA, checks can fail or error, and some runs may lack quality evidence altogether.
You must determine which pipelines need quarantine, replay or monitoring and replace the single green status with a defensible composite release gate.
Choose the pipelines to quarantine, replay or monitor and define the minimum evidence a future run must satisfy before publishing downstream data.
A successful orchestration state is not data certification and absence of a quality result is not a pass. Reduce checks to declared run-level controls before joining, and do not weight pipelines merely because they execute more checks.
Run the brief as a controlled assignment.
The workbench mounts only the listed source neighborhood and supplies neutral starter worksheets. It does not grade the conclusion or reveal the mechanisms planted in the larger assignment.
- Open the dedicated brief workspace and confirm the brief ID in the queue.
- Establish table grain, cutoff and control totals before joining or modeling.
- Make at least two distinct evidence moves and test a credible rival explanation.
- Leave the requested polished artifact, then export the workspace or portfolio package.
The source files are shared; the draft workspace is not. Changing to another full assignment drops brief mode deliberately.
Real tables. A deliberately bounded neighborhood.
These Parquet files already belong to Meridian’s public 96-table estate. Use the mounted schema.table names in SQL or Python; download links are provided for learners working in a local DuckDB environment.
GRAIN / One registered data pipeline.
CAUTION / Ownership, schedule, SLA and criticality are contracts rather than evidence that an execution was safe.
ANALYSIS CUTOFF / 20 Jan 2026 / 09:15 ET; runs completed through 31 Dec 2025
Questions to pressure-test—not steps to copy.
These prompts define the analytical territory without prescribing an order, technique or conclusion.
- 01
How often do successful runs also contain failed or errored checks, exceed SLA or lack quality evidence?
- 02
Which pipelines and quality dimensions show recurring rather than isolated problems?
- 03
How do conclusions change when calculated per check, per execution and per pipeline?
- 04
What evidence should be mandatory before a run can publish?
Leave work another analyst can review.
Artifact presence can be recorded; analytical quality remains a human judgment. A complete brief has evidence, reasoning and a decision—not merely executed code.
- 01Run-level control matrix
Combine orchestration, retry, SLA, row-count and quality states at a declared logical-run grain.
- 02Pipeline incident Pareto
Distinguish recurring, acute, missing-control and high-criticality problems without check-row weighting.
- 03Composite release gate
Define executable PASS, HOLD, REPLAY and REVIEW states with explicit missing-evidence behavior.
- 04Reliability handoff
Name the quarantine, replay and monitoring set and the evidence required to clear each action.
Review the reasoning after a real attempt.
The debrief does not contain an official answer. It identifies defensible analytical moves, common failure modes and questions a reviewer may use to challenge the handoff.
SPOILER-GATED REVIEW / REVEAL AFTER YOUR FIRST HANDOFF
DEBRIEF REVEALED / THIS MAY CHANGE HOW YOU APPROACH THE BRIEF
Orchestration, timeliness and data validity are separate dimensions. A useful release gate retains all three and treats missing control evidence explicitly.
Defensible approaches
- Resolve retries into a declared logical-run policy and reduce quality rows to run-level states before joining.
- Preserve executions with no quality evidence, evaluate runtime against each pipeline SLA and compare per-run with per-pipeline views.
- Combine criticality, persistence and control state in the action slate rather than ranking on raw check counts.
Common traps
- Treating SUCCEEDED as trusted or inner joining away runs with no quality results.
- Calculating a failure rate across check rows and overweighting pipelines that execute more checks.
- Counting retries or repeated incident references as independent business incidents without a declared rule.
Reviewer questions
- Could an orchestration-success run still publish unsafe data under the proposed gate?
- Which missing evidence produces a hold rather than an assumed pass?
- Does the triage order reflect business criticality as well as event frequency?
CHALLENGE A COLLEAGUE
Pass the Brief
The link shares this spoiler-free briefing. It never includes your work, identity, or browser progress.SUGGESTED NOTEThe pipeline was green. Its release evidence was not. I built the control gate in Green Run, Red Data from The Analyst.