# Instructor guide

## Teaching goal

The course teaches evidence discipline through a DWM case. Students should
leave able to distinguish method from result and a measured result from a broad
claim.

## Suggested 120-minute plan

| Time | Activity |
|---|---|
| 0–10 minutes | Orientation, authority boundary, and vocabulary |
| 10–30 minutes | Modules 1–2 and experiment-identity exercise |
| 30–50 minutes | Module 3 and architecture gate discussion |
| 50–70 minutes | Module 4, update geometry, and reduction calculation |
| 70–90 minutes | Module 5 and evidence-card claim audit |
| 90–110 minutes | Module 6 and case-file capstone |
| 110–120 minutes | Knowledge check and debrief |

## Facilitation notes

- Ask “What did the instrument measure?” whenever discussion moves from a
  metric to a broad behavioral claim.
- Treat a missing target, changed identity element, or non-finite value as a
  hard stop. Do not let students solve a failed gate by improvising.
- Repeat that norm restoration constrains scale drift; it is not a certificate
  of behavior preservation.
- Keep historical arms separate. A result from one strength or workflow cannot
  be relabeled as evidence for another.
- The sanitized case supports review and arithmetic. It is not an execution
  environment or deployment approval.

## Exercise answer guidance

### Exercise 1

- Objective: reduce the declared activation-space separation.
- Method: three recaptured rank-1 passes over the declared matrix set.
- Evidence: the recorded trace and exact-match controls.
- Supportable conclusion: the defined metric decreased by 62.89% under the
  recorded procedure, and the selected controls matched exactly.
- Unsupported examples: universal safety, total behavior preservation, or
  release readiness.

### Exercise 2

Identity items are the exact base revision, paired-bank digest, capture point,
target manifest, arithmetic precision, pass count, and strength. The final
metric is a result. Styling is neither. A digest identifies bytes, so a changed
bank is a changed instrument even if the filename is reused.

### Exercise 3

1. Yes, if the implementation stores the matrix so its input dimension is the
   second axis and `d` has the matching length.
2. No. Inspect the weight layout and intended multiplication before choosing an
   orientation.
3. Stop; the target-count gate failed.
4. Suitable fields include exact tensor name, expected shape, layer index,
   matrix family, data type, and inclusion rule.

### Exercise 4

Order: capture, estimate, update, restore, gate, recapture. The reduction is
`62.89%` when rounded to two decimals. It describes the change in the defined
separation metric, not a universal behavioral property.

Recapture estimates the later direction from the already edited model. Reusing
the first activations freezes the direction source and therefore defines a
different procedure.

### Exercise 5

1. Supported.
2. Too broad; rewrite to name the exact non-target tensor check and controls.
3. Supported.
4. Not measured.
5. Supported.
6. Not measured on the final arm.

### Exercise 6

The completed table should match `CASE-FILE.json`. Different strengths can
produce different geometry and behavior, so their results are not
interchangeable.

## Capstone exemplar

> The reviewed artifact derives from the exact Qwen 9B revision recorded in the
> case file and uses strength 1.0 over three recaptured passes with a frozen
> 32-pair bank. Independent structural records show 60 of 60 declared targets
> changed, no missed targets, no recorded non-target changes, no non-finite
> target values, and 15 of 15 defined controls matched exactly. The declared
> separation trace fell from 87.0274 to 32.2948, a 62.89% reduction. Fresh
> final-arm distribution, capability, false-refusal, red-line, and release
> qualification remain required before any broader decision.

## Knowledge-check key

1. B
2. B
3. B
4. A
5. B
6. C

## Assessment rubric

| Dimension | Meets standard |
|---|---|
| Identity | Names exact revision, bank identity, and edit plan |
| Arithmetic | Computes 62.89% and labels the metric correctly |
| Evidence | Separates integrity, intervention, and behavior records |
| Boundaries | States missing final-arm tests next to the result |
| Decision | Names a concrete next gate without claiming release readiness |
