# Verification Report

Audited 2026-10-06 against the source course, the downloaded media, and the live deployment.
Every number below comes from a script run in this pass, not from an earlier claim.

## 1. Deployment matches the local build

| Check | Result |
|---|---|
| `https://gcc-cbyx-archive.pages.dev/` vs `site/index.html` | byte-identical (sha256 of both files compared at deploy time; the digest is not printed here because this report is embedded in that file, so printing it would be stale by construction) |
| `/lessons.json` live vs local | byte-identical (sha256 `ec5e6f66…`) |
| `/templates/*` (MANIFEST, 13 files) | HTTP 200, byte-identical to local |
| `/exports/*` (course.md, kernels, verification, transcripts) | HTTP 200, byte-identical to local |
| Transfer size | html 64.7 KB compressed, lessons.json 350 KB compressed (brotli) |
| Deployment branch | production (`--branch=main`, matches the project's production branch) |

## 2. The scrape matches the live course page

Compared `gcc-pack/structure.json` against the course page captured from the live site
(`gcc-course-map.txt`, 2026-09-16): **9 modules and 46 lessons, zero mismatches** in module
names, lesson titles, or ordering.

Re-crawled the live site on 2026-10-06 with a real login (all 46 lesson pages, HTTP 200):

- Outline is identical to the archive: 9 modules, 46 lessons, same titles, same order,
  same URLs. 44 lessons have a video, 2 do not (Support, Resources).
- **8 of the 44 lesson-to-video bindings in the archive were wrong.** Each of those 8
  lessons pointed at a *different* lesson's video, so the site was showing another
  lesson's transcript and summary under the wrong title.

| Lesson | archived (wrong) | live and correct |
|---|---|---|
| What You Need to Know about Internships in Germany | 662302872 *Foundation Learning Tool #4* | 863060446 *What You Need to Know About Internships in Germany* |
| Overview of the Internship Search Process | 662303012 *Using Professional Learning Tools* | 862938711 *Internship search process* |
| The Internships in Germany Roadmap | 662307174 *While Job Searching* | 862942738 *Roadmap* |
| The International Career Management System | 662307959 *On the Job* | 862938725 *ICMS* |
| Timing and Time Plan | 662310708 *What to Avoid* | 863060406 *Timing* |
| Leveraging LinkedIn to Establish a Germany Niche | 662312794 *How to Set Up Your 30-Day Framework* | 862188336 *LinkedIn Quick Start Training* |
| Cover Letters, Motivation Letters and Emails | 862532275 *Dan Baxter CBYX Interview* | 864692415 *cover letters* |
| Contracts and Salary Negotiation | 853991974 *Jamauri Bryan CBYX Interview* | 862693316 *Module 5 part 2* |

  The video titles above come from Vimeo's own oEmbed endpoint, not from an assumption:
  the live ID on every one of these lessons carries a title that matches that lesson, while
  the archived ID carries a title belonging to another lesson. The bindings were wrong from
  the first capture on 2026-09-16 (the same 8 errors are in `gcc-videos.json`, the original
  session's video map), so the error propagated into `meta.json`, `structure.json`, and from
  there into the transcripts, summaries, and the site.

- Fix applied in this pass: `meta.json` and `structure.json` rewritten from the 2026-10-06
  crawl. After the fix: **0 mismatches against the live site, 44 unique Vimeo IDs for 44
  videos, 0 duplicates.**
- The 8 videos that were never downloaded are now archived locally (audio, durations match
  Vimeo's own listing); their transcripts and summaries are being generated.

**Correction to an earlier claim:** the live page has **no video** on the *Support* and
*Resources* lessons. An earlier version of this report said those two lessons embed Vimeo
`862188336` and that the download failed with a 403. Both halves of that were wrong:
`862188336` is the *Leveraging LinkedIn* lesson's video (now archived), and Support /
Resources genuinely have no video, which the site now says.

## 3. Transcripts faithfully reproduce the audio

Re-transcribed 20 random 30-second windows (from 12 different videos, early and late in each)
with the same model family (`large-v3-turbo`) and compared word-by-word against the stored
transcripts:

| Metric | Result |
|---|---|
| Mean sequence similarity (token order + wording) | **0.882** |
| Mean share of transcript words present in the fresh transcription | **0.948** |
| Windows checked | 20 |

Residual differences are punctuation, contractions and paraphrase between ASR runs, not
missing or invented speech.

### Defects found and fixed in this pass

Five transcripts had whisper "hallucinated" tails past the end of the audio. Trimmed to the
real duration (plus a manual cut where the speech had clearly stopped):

| Video | Was | Now | Problem |
|---|---|---|---|
| `662297870` | 95 lines | 85 | invented text from [04:54], incl. stray CJK characters; audio ends 4:55 |
| `662302872` | 32 lines | 16 | "Thank you" loop + counting ("30 40 50") past 1:28 |
| `662311691` | 31 lines | 22 | "Bye." repeated past 1:12 |
| `865292040` | 817 lines | 809 | "Bye." loop past 43:21 |
| `662300189` | 76 lines | 75 | duplicated sign-off past 2:59 |

Originals kept as `*.txt.bak`. Transcripts are monotonic, and no line now runs past its audio.

## 4. Summary citations land where they claim

| Check | Result |
|---|---|
| Citation endpoints in the summaries (ranges count both ends) | **890** |
| Within 3 s of an actual transcript timestamp | **881** (99.0%) |
| Inside the utterance they point at | **890 (100%)** |
| Pointing past the end of their audio | **0** |
| Why strict is not 100% | several transcripts have 30-second segments, so a citation at e.g. 00:38 sits inside the utterance that opens at 00:30 but not within 3 s of it |


### Per-lesson citation audit (re-run 2026-10-06)

`strict` = endpoint within 3s of a transcript timestamp. `span` = endpoint falls inside
the utterance that timestamp opens (the honest test for transcripts with long segments).

| video | cites | strict | span | audio end |
|---|---|---|---|---|
| 1157426586 | 64 | 64 | 64 | 1839s |
| 662280404 | 8 | 8 | 8 | 47s |
| 662297591 | 14 | 14 | 14 | 83s |
| 662297870 | 23 | 23 | 23 | 295s |
| 662298334 | 29 | 29 | 29 | 561s |
| 662300021 | 18 | 18 | 18 | 137s |
| 662300189 | 20 | 20 | 20 | 179s |
| 662300414 | 16 | 16 | 16 | 99s |
| 662300580 | 22 | 22 | 22 | 240s |
| 662301742 | 20 | 20 | 20 | 228s |
| 662302872 | 19 | 19 | 19 | 88s |
| 662303012 | 16 | 16 | 16 | 98s |
| 662307174 | 24 | 24 | 24 | 251s |
| 662307959 | 20 | 18 | 20 | 232s |
| 662308798 | 19 | 15 | 19 | 159s |
| 662310708 | 20 | 20 | 20 | 179s |
| 662311691 | 15 | 15 | 15 | 72s |
| 662312396 | 20 | 17 | 20 | 288s |
| 662312794 | 26 | 26 | 26 | 326s |
| 662313488 | 21 | 21 | 21 | 189s |
| 775823994 | 19 | 19 | 19 | 390s |
| 853991974 | 15 | 15 | 15 | 1609s |
| 854039627 | 14 | 14 | 14 | 1018s |
| 862123445 | 38 | 38 | 38 | 2333s |
| 862532275 | 20 | 20 | 20 | 2154s |
| 862649036 | 18 | 18 | 18 | 576s |
| 862649743 | 13 | 13 | 13 | 1131s |
| 862691465 | 35 | 35 | 35 | 1229s |
| 862720732 | 20 | 20 | 20 | 1708s |
| 862954055 | 21 | 21 | 21 | 839s |
| 862969032 | 34 | 34 | 34 | 1336s |
| 863060318 | 25 | 25 | 25 | 632s |
| 864696047 | 11 | 11 | 11 | 402s |
| 864705843 | 15 | 15 | 15 | 281s |
| 865292040 | 145 | 145 | 145 | 2601s |
| 865333620 | 13 | 13 | 13 | 372s |
| **total** | **890** | **881** | **890** | |

## 5. Numbers in the summaries are in the transcripts

Extracted every numeric claim from the 36 summaries (2091 tokens after normalising
"two months" = "2 months" etc.) and required each to appear in the matching transcript:
**all 2091 found.** Sample: 99.6% / 35-36% / 60% Mittelstand figures traced to
`862969032` [18:39-19:13]; "10 hours", "20 hours" to `853991974` [07:56], [08:22].

## 6. Claims spot-checked by hand (paraphrase check)

A word-overlap heuristic flagged 98 of 517 cited claim bullets as "mostly reworded". Five were
read back against the transcript; all five were accurate paraphrases:

| Claim | Transcript |
|---|---|
| notifications trial for 1-2 weeks then prune | [06:13] "first week or two... keep for good... deactivate" |
| 5-10 informational invites -> 1-2 replies | [04:19] "aim for five to 10... might get one or two back" |
| requested Arbeitszeugnis from both roles | [18:09] "get like, a testimonial... in English and German" |
| keyword stuffing annoys the human reader | [10:29] "may be counterproductive... that's going to be annoying" |
| LiveLingua for role-playing hard conversations | [00:16] "point you in the direction of life lingua" |

## 7. What is NOT verified

- **Word-level accuracy of all 36 transcripts** - 20 windows were re-transcribed, not all
  5,171 lines. The 8 transcripts produced in this pass are not yet in that sample.
- **The 8 newly archived videos** - audio downloaded and durations checked against Vimeo;
  they have not been re-downloaded independently, and their transcripts/summaries are new.
- **The two lessons with no video** (Support, Resources) have no audio to check.
- Claims not expressible as a number and not in the hand-checked sample above.

## 8. The published site was exercised in a browser

Headless Chrome against `https://gcc-cbyx-archive.pages.dev/` on 2026-10-06: 9 modules and
46 lessons in the sidebar, takeaways box and 8 citations on a sample lesson, transcript pane
(13 lines on that lesson), Copy / Export with 6 cards ("copy entire course" = 189,917
characters), kernels / templates / verification views, search (Bewerbung -> 4 hits), `/`
focuses the search box, mobile drawer opens, theme cycles system -> dark -> light -> system.
Zero console errors, zero failed requests.

Three defects found and fixed in this pass:

| Defect | Symptom | Fix |
|---|---|---|
| the mobile backdrop div took part in the desktop grid | sidebar sat in the 1140px column and the article was squeezed into 300px below it; live since the first deploy | `.side-backdrop{display:none}` outside the mobile media query |
| the theme button could never reach dark | the next state was derived from the OS preference, so a light-OS visitor was stuck on system | explicit system -> dark -> light -> system cycle |
| markdown tables rendered twice | table HTML followed by the same rows as literal pipe paragraphs | the converter assigned to a `for` loop variable, which Python ignores; consumed rows are now skipped |

## Result

Structure matches the live course page (9 modules, 46 lessons, 44 videos, 0 mismatches
after fixing 8 lesson-to-video bindings that had been wrong since the first capture);
deployment matches the local build byte-for-byte; every citation resolves to real speech
inside its video; every numeric claim traces to the transcript; five transcripts were
corrected where whisper had invented or looped text; and three rendering defects in the
site itself (grid, theme toggle, table converter) were found and fixed. Two lessons
(Support, Resources) have no video on the live page and say so; 8 previously missing
videos are now archived and their transcripts are in progress.
