Turn the noisy, multi-format transcript of an AI agent's coding session into a single self-contained HTML walkthrough that teaches a teammate what was built, why, and how — without making them read the raw logs.
Walkthrough · The walkthrough product as it stands now: the transcript->HTML pipeline, the output schema + editorial model, the two architectures, the quality gate, and the renderer/viewer

A six-stage transcript-to-HTML pipeline that emits one offline, evidence-backed, dual-view walkthrough from any of three agent session providers.

Turn the noisy, multi-format transcript of an AI agent's coding session into a single self-contained HTML walkthrough that teaches a teammate what was built, why, and how — without making them read the raw logs.

A 75-second tour — from transcript chaos to one offline walkthrough you can trust.
  • Ingest normalizes Codex, Claude, and OpenCode logs into one event model; reduce projects + chunks them deterministically; summarize fans out per-chunk LLM subagents into a draft; author reshapes the draft into a Descent or Journey; verify+render gate-checks and emits the HTML.
  • The reduce stage is deterministic (no LLM): projection achieves a measured 61% byte reduction (63MB->25MB, 217->85 chunks) on the calibration dataset.
  • Trust is enforced by two validators backed by a ~245-test suite: a structural pipeline contract checker and an editorial quality gate that refuses to render a failing walkthrough without --allow-draft.
  • The single HTML file inlines every image, the ~2.4 MB LikeC4 (architecture-diagram) bundle, and all JS — zero network dependencies — and enriches itself at load with prose links, glossary tooltips, and diff counts.
  • A deterministic projection stage cuts ~60% of bytes before any LLM runs, so each chunk carries more signal.
  • A quality gate validates honesty and source-ref integrity and binds at render time — a failing walkthrough is a draft.
  • The output is one offline HTML file with a dual-view viewer, a hover glossary, and an interactive architecture diagram.
01 The destination: a six-stage transcript-to-HTML pipeline What exists now is a six-stage pipeline — ingest, reduce, summarize, author, verify, render — that turns any agent session log into one self-contained HTML walkthrough, which is... 02 Ingest: three providers folded into one event model The front of the pipeline finds sessions across three providers, strips binary blobs to protect the context window, and normalizes everything into 15 event kinds — which is why... 2 decisions 03 The output contract: an altitude ladder and two architectures What lands on the page is an altitude ladder per step — title, takeaway, intent, narrative, collapsed proof — assembled as one of two measured architectures, the Descent or the... 2 decisions 04 How much can I trust it: two layered validators Trust rests on two layered validators backed by a ~245-test suite — a structural contract checker plus an editorial quality gate that binds at render time — which is why a... 2 decisions 05 Why reduce before reasoning: deterministic projection A no-LLM projection stage drops zero-signal events and compresses successful tool output for a large measured byte cut on the calibration dataset — which is why each chunk reaches... 2 decisions 06 Why fan-out summarization with cache-by-hash Per-chunk LLM subagents emit JSON summaries keyed by chunk_id plus the chunk's sha256, and merge_summaries folds them into a mechanical draft — which is why a re-chunked transcript... 2 decisions 07 Why one offline file with optional viewer chrome render_html.py emits a single self-contained HTML file that inlines every image, the ~2.4 MB LikeC4 (architecture-diagram) bundle, and all JS, then enriches itself at load — which... 2 decisions 08 Where to start: SKILL.md, the script table, and the doctor A teammate starts at SKILL.md, runs check_setup.py to confirm the environment, and uses the script-to-skill table to find the one stage to change — which is why extending the...

System

The system today

Architecture overview diagram Architecture overview diagram
The end-to-end pipeline. Click into any cluster (ingest, reduce, summarize, author, verify+render) for the next level down.
How it works today
Ingest (discover + strip + normalize) discover_sessions finds sessions for three providers, strip_binary removes base64/oversized fields, and three normalizers fold each provider's native format into one normalized event model (15 event kinds). Step 2 →
Reduce (project + chunk + cards) Deterministic, no-LLM: project_events drops noise and compresses successful tool output (~60%), chunk_events splits into byte-bounded chunks that never break a tool_use/tool_result pair, extract_session_cards emits a ~2KB overview per session. Step 2 →Step 5 →
Summarize (chunk -> session -> draft) Per-chunk LLM subagents emit cache-by-hash JSON summaries; merge_summaries folds them into a mechanical 1:1 chunk-to-step draft-walkthrough.json. Step 6 →
Author (editorial assembly) An editorial agent reshapes the draft into 8-12 teaching steps shaped as a Descent or Journey, chosen from the reader frame by a measured selection rule. Step 3 →Step 7 →
Schema + editorial model walkthrough.json is an altitude ladder (title -> takeaway -> intent -> narrative -> collapsed proof) with confidence-tagged claims kept separate from verbatim, grounded-by-construction evidence. Step 3 →Step 7 →
Verify + Render (gate + viewer) validate_pipeline checks structural contracts; validate_walkthrough_quality is the editorial gate that binds at render time; render_html emits the self-contained dual-view HTML viewer. Step 4 →Step 8 →
Current constraints
  • Reduce-stage byte reduction is measured on a single calibration dataset cited in SKILL.md (the 61% figure stated above); it is not a guaranteed ratio across arbitrary inputs. Step 5 →
  • Trust rests on a ~245-test suite across 15 files; the editorial quality gate has 43 tests and the structural pipeline validator has 5 (only validate_normalized is unit-tested — validate_projected and validate_chunks have no dedicated test). Step 4 →
  • The whole pipeline reads/writes a single hardcoded out/ namespace; a second walkthrough silently clobbers the first, so an extra run must be isolated in out/<name>/ for every stage.
  • Screenshot bytes survive only in normalized.jsonl: normalize_claude carries media.data_b64 forward but project_events strips it from the chunking stream, so render_html.resolve_media re-reads it from normalized.jsonl — a render without --normalized silently drops screenshots from a media-bearing walkthrough.
  • validate_pipeline's turn_index monotonic check is a known false-positive on /compact-ed or resumed sessions; the documented --allow-turn-regression fix is an unimplemented TODO and its named symbol (_validate_stream_file) is stale.
  • render_html.py is the only stage with third-party deps (jinja2, pygments) via an inline PEP 723 block; it must be invoked as `uv run scripts/render_html.py`, not `uv run python3 scripts/render_html.py`, or the block is ignored. Step 8 →
  • The architecture-selection rule defaults to the Descent on ambiguity because blind judge panels measured it sweeping destination clarity; that is an empirical fault line, not a universal preference. Step 3 →

Step 1

The destination: a six-stage transcript-to-HTML pipeline

What exists now is a six-stage pipeline — ingest, reduce, summarize, author, verify, render — that turns any agent session log into one self-contained HTML walkthrough, which is why projection, chunking, and summarization all read one common event model instead of three provider formats.

Before any component, a teammate needs the whole shape: what the system takes in, the stages it flows through, and what it emits. This first tour step anchors on the architecture diagram on the synthesized system section.

The pipeline's operational procedure is documented end to end in SKILL.md's numbered Workflow (sections 1-8) as an ordered sequence: discover -> strip_binary -> per-provider normalizer -> project -> chunk -> summarize -> merge -> editorial assembly -> quality gate -> render.

Normalization gives projection, chunking, and summarization one common schema — the normalized event model — so those consumers serve Codex, Claude, and OpenCode without provider-specific branches; the later stages (merge, render) consume the summaries and walkthrough.json those earlier stages produce.

The final artifact's primary output is walkthrough.json — a structured, evidence-backed narrative — which render_html.py turns into a single HTML file.

inferredThe clusters in the architecture diagram (ingest, reduce, summarize, author, verify+render) map one-to-one onto the script families that implement them, so the diagram is a faithful index into the codebase.

Step 2

Ingest: three providers folded into one event model

The front of the pipeline finds sessions across three providers, strips binary blobs to protect the context window, and normalizes everything into 15 event kinds — which is why downstream stages never see a provider's idiosyncratic format.

Component one of the destination tour: how heterogeneous Codex CLI, Claude Code, and OpenCode logs become uniform, LLM-safe input.

discover_sessions.py finds sessions for three providers with three mechanisms: Codex via a recursive glob of rollout-*.jsonl, Claude via */*.jsonl, and OpenCode via a SQLite query.

strip_binary.py's base64 detector replaces any string longer than 1000 chars that fully matches the base64 alphabet with a [BASE64: N bytes, source: path:line] marker, recording the byte count and source so nothing is silently lost.

Claude Code emits no explicit file-change events, so normalize_claude.py fabricates a synthetic diff for each Edit — emitting all old lines as '-' then all new lines as '+' under a single @@ hunk header — to give downstream a uniform file_change kind.

The normalized contract is enforced by validate_pipeline.py: every event must carry the required envelope fields and the provider must be one of codex, claude, or opencode.

◆ Decision

Normalize three heterogeneous provider formats into one common JSONL event model.

Chunking, summarization, and assembly must be provider-agnostic; a single schema lets one set of consumers serve all three providers.

  • Let each downstream stage understand each provider's native format directly.
◆ Decision

Strip base64/binary content as a separate pre-normalization pass with provenance markers.

Base64 screenshots and blobs blow up the LLM context window; a single recursive strip pass keeps a byte-count + source marker so the data is accounted for and runs once regardless of provider.

  • Drop or truncate fields ad hoc inside each normalizer
  • Pass full base64 into the LLM context
Step 3

The output contract: an altitude ladder and two architectures

What lands on the page is an altitude ladder per step — title, takeaway, intent, narrative, collapsed proof — assembled as one of two measured architectures, the Descent or the Journey, which is why a reader can stop at any depth and still get a coherent answer.

Component two of the destination tour: the schema, editorial model, and the architecture choice that govern what a walkthrough looks like and how its steps are shaped. The skim test — reading the takeaway lines alone as a complete summary — is the core editorial constraint both architectures answer to.

Each step is rendered as a five-rung descent — title -> takeaway -> intent -> narrative (claims + decisions + gotchas) -> proof (evidence in a collapsed details) — so a reader can stop at whatever depth they need.

Claims carry exactly three confidence levels — grounded (rendered plain), inferred (subtle indicator), speculative (visible badge) — while evidence fields carry no confidence because they are grounded by construction.

Both architectures share one schema, gate, altitude ladder, and evidence rules; only the step structure differs, and that structure is chosen from the reader frame before any steps are written.

The Descent answers the reader's five questions in order — what exists now, how much can I trust it, why is it shaped this way, what fought back, where do I start — and the takeaway sequence read alone is those answers; this walkthrough is itself a Descent, here locked to its End State view, so its 'what fought back' answer lives in the current-constraints reference rather than its own step.

◆ Decision

Separate narrative (confidence-tagged claims the model authors) from evidence (deterministic, verbatim, grounded-by-construction artifacts).

A reader must be able to verify the artifact against the raw transcript; paraphrased code is unfalsifiable, so diff_hunks and commands are verbatim or absent — if only prose survives, it is demoted to a claim with a source_ref.

  • Let the model write everything, including diff-like before/after summaries, with a confidence level.
◆ Decision

Choose the step architecture from meta.purpose and meta.audience up front, defaulting to the Descent on ambiguity.

Blind judge panels measured the question-led Descent sweeping destination clarity for handoff readers, and a destination-first artifact degrades more gracefully for the wrong reader than a chronology does.

  • Default to the Journey
  • Decide the structure late, after drafting steps
  • Treat the two architectures' guidance as co-equal
Step 4

How much can I trust it: two layered validators

Trust rests on two layered validators backed by a ~245-test suite — a structural contract checker plus an editorial quality gate that binds at render time — which is why a walkthrough that fails the gate cannot become a shared HTML file without an explicit --allow-draft.

The proof-ledger step: in one place, what is verified, by what, and where the coverage thins out — including the honesty checks that make confidence labels real.

The editorial gate hard-errors any step that lacks at least one grounded claim carrying a non-empty source_refs list — _has_grounded_source_ref returns True only on a grounded confidence with refs present.

The gate binds at render time: render_html.py re-runs validate_walkthrough and exits 1 (printing QUALITY GATE ERROR lines) unless --allow-draft is passed.

The numeric trust thresholds are explicit constants: source-ref spans cap at 200 lines, a shared range cited by 3+ claims is flagged, an all-grounded confidence monoculture warns at 20+ claims, and the glossary caps at 50 entries / 300-char definitions.

The structural validator covers different ground: it checks turn indexing per conversational stream, requiring the meta event at turn_index 0 and the first user_message at turn_index 1 — proving the data is sound before the LLM stage spends tokens.

The mechanical gate is the executable analog of the human TTG judge rubric, whose five criteria (destination clarity, story coherence, altitude correctness, signal density, evidence trust) are scored 1-5 by a blind panel.

Coverage is uneven by design: the editorial gate has 43 test functions (counted via grep -c 'def test_' in test_validate_walkthrough_quality.py) but the structural validator's 5 tests import and exercise only validate_normalized (test_validate_pipeline.py:8), leaving validate_projected and validate_chunks without dedicated unit tests.

◆ Decision

Split validation into a structural pipeline-contract checker and a separate editorial finished-walkthrough gate.

Pipeline validation proves the event data is structurally sound; the gate catches valid-JSON-but-poor-reading drafts — different failure classes at different stages need different checks.

  • A single validator that checks both data soundness and reading quality.
◆ Decision

Make warnings advise but never block; only a curated set of reader-trust-destroying defects are errors.

Most quality smells are editing instructions the viewer can survive (it clamps over-tall steps), so blocking on them would make the gate un-shippable.

  • Treat every smell (provenance, glossary, height, diagram) as a hard failure.
▸Evidence2 files · 3 cmds
Commands
python3 scripts/validate_walkthrough_quality.py --input out/walkthrough.json --max-steps 12
The documented gate invocation; exits non-zero on any error, zero when only warnings remain.
grep -cE '^\s*def test_' tests/test_validate_walkthrough_quality.py
43 — the editorial gate's test count.
grep -cE '^\s*def test_' tests/test_validate_pipeline.py
5 — the structural validator's test count.
Step 5

Why reduce before reasoning: deterministic projection

A no-LLM projection stage drops zero-signal events and compresses successful tool output for a large measured byte cut on the calibration dataset — which is why each chunk reaches the summarizer carrying roughly 3x more reasoning per byte.

Decision one: the reduce stage exists because Codex sessions emit huge zero-signal events that would otherwise dominate the LLM budget. The same stage also extracts a ~2KB session card per session for editorial context. Each claim here is a decision with its rejected alternative.

Projection drops exactly three event kinds — file_snapshot, turn_context, and compaction — defined in the module-level DROP_KINDS set.

On the calibration dataset SKILL.md records projection cutting 63MB/44K events to 25MB/38K events — a 61% byte reduction — and chunk count from 217 to 85.

Chunking keeps tool pairs intact by defining an atomic group as a tool_use plus its matching tool_result plus adjacent file_change/command — but only when they share the same session AND agent, so a tool_use from another stream does not close the current group.

◆ Decision

Insert a separate projection stage between normalize and chunk rather than compressing inline.

Codex sessions emit huge zero-signal events (file_snapshot dumps, ~12KB turn_context per turn) that blow up chunk byte budgets; projecting them out first triples reasoning density while keeping normalized.jsonl as the full-fidelity reference.

  • Normalize directly into chunks with no noise-reduction layer (the original --no-project path).
◆ Decision

Compress non-error tool_results to byte/line counts plus a head line, but keep error tool_results at full fidelity.

Successful tool output is bulky and low-signal, but error output is exactly what the summarizer needs to explain what fought back; the asymmetry sheds volume while preserving debugging signal.

  • Compress all tool_results uniformly
  • Keep all tool_results verbatim
Step 6

Why fan-out summarization with cache-by-hash

Per-chunk LLM subagents emit JSON summaries keyed by chunk_id plus the chunk's sha256, and merge_summaries folds them into a mechanical draft — which is why a re-chunked transcript never reuses a stale summary and the editorial agent always starts from a clean 1:1 chunk-to-step base.

Decision two: the summarize stage is parallel and content-addressed on purpose. Each claim pairs the design with the constraint that forced it.

find_summary resolves a summary only by chunk_id plus the first 8 hex chars of the chunk's sha256, returning None when that exact file is absent — there is no fuzzy prefix fallback to an older summary.

merge_summaries produces a deterministic 1:1 chunk-to-step draft tagged scope 'draft — needs editorial reshaping', leaving grouping, ordering, and compression to the editorial agent.

inferredStep 5 fan-out is provider-asymmetric — Claude Code spawns chunk-summary subagents in parallel while Codex processes chunks sequentially — so the same logical stage runs differently per provider.

◆ Decision

Cache summaries by chunk_id + full sha256 and reuse only on an exact match.

Chunk content changes when upstream normalization or projection changes; reusing a summary keyed only by chunk_id would silently describe content that no longer exists.

  • Reuse any older summary whose filename shares the chunk_id prefix.
◆ Decision

Have merge_summaries emit a mechanical draft, not the final artifact.

Grouping, ordering, and omission are editorial judgment driven by the reader frame; a script can't make those calls, so it stays deterministic and tags the draft for reshaping.

  • Have the script itself decide step grouping, ordering, and compression.
Step 7

Why one offline file with optional viewer chrome

render_html.py emits a single self-contained HTML file that inlines every image, the ~2.4 MB LikeC4 (architecture-diagram) bundle, and all JS, then enriches itself at load — which is why the artifact has zero network dependencies; separately, its meta.ui chrome flags default ON and coerce missing values so older walkthroughs render unchanged.

Decision three: the renderer's single-file, enrich-at-load design. The cover of this very walkthrough embeds a HyperFrames overview tour to displace summary prose — media_mode=none refers to screenshot extraction, which this artifact doesn't use; the video is a separate overview.video embed. Each claim pairs the choice with the constraint behind it.

Files touched

render_html.py is the only pipeline script with third-party dependencies (jinja2, pygments) declared in an inline PEP 723 block, and its docstring warns that invoking it as `uv run python3 scripts/render_html.py` ignores the block.

meta.ui's four chrome flags (view_switcher, present_mode, stats, confidence_legend) all default ON via resolve_ui_flags, which coerces any missing or non-bool value to its default so existing walkthroughs render unchanged.

Prose links are applied client-side in three layers before the glossary pass so links never double-wrap, and the glossary annotates matching terms into hover tooltips after load — neither pass injects into code, diffs, commands, or controls.

◆ Decision

Make all four meta.ui chrome flags default ON and coerce malformed values to their default.

Backward compatibility — every walkthrough.json produced before meta.ui existed has no ui block, so default-ON coercion is what keeps those existing artifacts byte-identical in render, and a malformed flag degrades gracefully instead of throwing.

  • Require walkthroughs to opt in to chrome features
  • Hard-fail on malformed meta.ui
◆ Decision

Inline the LikeC4 bundle by default rather than referencing it as a sidecar file.

The single-file/offline guarantee for the artifact; embed:false is offered as an explicit opt-out when several walkthroughs share one bundle.

  • Keep the bundle as a sidecar <script src> referenced relative to the HTML.
Step 8

Where to start: SKILL.md, the script table, and the doctor

A teammate starts at SKILL.md, runs check_setup.py to confirm the environment, and uses the script-to-skill table to find the one stage to change — which is why extending the pipeline is a matter of editing one script and re-running its tests, not learning the whole flow at once.

The orientation step: the entry-point files, the readiness check, and how to extend the system. Tagged end-state so it lands as a reference, not a chronology beat.

Files touched

SKILL.md is the operational entry point: it documents every stage's command invocation and ordering, and its script table maps each script to its skill name, role, and input/output.

The first-run readiness check is check_setup.py (skill walkthrough-doctor), which reports environment readiness and flags missing optional tooling like LikeC4 before a first walkthrough.

inferredEach stage is an independent script with its own test file, so extending the pipeline means editing one script (e.g. scripts/project_events.py or scripts/render_html.py) against the contract the next stage's validator enforces.

▸Evidence1 file · 1 cmd
Commands
python3 scripts/check_setup.py
First-run environment doctor; reports readiness and flags missing optional tooling (LikeC4, HyperFrames).