# Course transcript

## Orientation: read this as an evidence course

Domain Weight Modification (DWM) is the method under study, but the
transferable skill is disciplined inference. At every stage, ask:

- What was fixed?
- What changed?
- What was measured?
- Which controls held?
- What conclusion is supported?
- Where does the evidence stop?

This material contains no model weights, raw paired bank, activation tensors,
or execution code. It does not authorize work on systems or artifacts you do
not own or have explicit permission to evaluate.

## Module 1 — Scope and vocabulary

DWM is a targeted model-weight intervention guided by measured activation
differences. Its purpose is to change a defined geometric relationship in
selected weights while keeping the procedure bounded and auditable.

Do not collapse these distinct methods:

- **DWM** measures a direction from paired observations and applies a
  constrained update to selected matrices.
- **Fine-tuning** optimizes parameters against a training objective, usually
  through gradient-based learning over a dataset.
- **Ablation** removes, disables, or zeros a component to test its causal role.
- **Prompt steering** changes runtime context or input without changing the
  stored weights.

Keep four layers separate. The **objective** says what relationship you intend
to change. The **method** says how. The **evidence** records what the instruments
observed. The **conclusion** must remain inside that evidence boundary.

If a geometric separation metric decreases, the supportable claim is that the
defined metric decreased under the recorded procedure. That observation alone
does not establish universal safety, complete capability preservation, or
release readiness.

## Module 2 — Measurement instrument

A result cannot be reproduced from a model name alone. The experiment identity
is a bundle: exact base revision, tokenizer behavior, paired bank, prompt
construction, capture point, target matrices, numerical precision, algorithm
parameters, and evaluation rules.

A paired bank presents two defined variants for each concept under comparison.
In a prefill-only procedure, activations are captured from the supplied token
sequence without sampling a continuation. Pair differences contribute to a
direction estimate:

```text
paired item -> tokenize -> prefill -> capture representation
            -> compute pair difference -> aggregate direction
```

Once accepted, the bank is immutable for that experiment. Record its SHA-256
digest and reject silent changes. If a pair changes during the procedure, the
measuring instrument changed too.

The minimum identity record includes:

- exact model and source revision;
- tokenizer and chat-template behavior;
- bank row count, pair count, and SHA-256 digest;
- capture point and token-position rule;
- selected layers and matrix families;
- pass count, strength parameter, precision, and norm policy; and
- gate thresholds and independent verification procedure.

If any approved identity element does not match, stop. A different identity is
a different experiment.

## Module 3 — Architecture and edit surface

The direction lives in residual-stream space. An edit affects a matrix only
when the matrix consumes a compatible representation, so architecture
inspection is part of the method.

```text
tokens -> residual state -> projection -> updated residual
```

For a matrix `W` with columns in the residual dimension and a unit direction
`d`, a rank-1 projection can remove the component of each row aligned with `d`:

```text
W' = W - alpha (W d) d^T
```

`alpha` is the edit strength. The exact orientation must be verified against
the implementation’s weight layout. A target manifest should enumerate exact
tensors, expected shapes, and expected count.

Architecture gates ask:

1. Does the captured representation dimension match the matrix input
   dimension?
2. Is every expected target present exactly once?
3. Are excluded tensors demonstrably unchanged?
4. Does serialization preserve names, shapes, and data types?

A mismatch is a failed gate, not a reason to guess.

## Module 4 — Iterative rank-1 editing

The runbook pattern is iterative. Each approved pass follows six stages:

1. **Capture:** run the frozen paired bank in prefill mode and collect the
   declared representations.
2. **Estimate:** aggregate normalized pair differences into a unit direction.
3. **Update:** apply the bounded rank-1 projection to every tensor in the
   target manifest.
4. **Restore:** rescale each matrix to its own pre-edit Frobenius norm.
5. **Gate:** check target count, non-target identity, finite values, controls,
   and recorded metrics.
6. **Recapture:** if another pass is approved, measure fresh activations from
   the newly edited model.

The per-matrix norm restoration is:

```text
W'' = W' * (FrobeniusNorm(W) / FrobeniusNorm(W'))
```

Norm restoration constrains scale drift. It does not prove that all behavior is
preserved, which is why the workflow also needs controls and broader evaluation.

Relative reduction is computed as:

```text
100 * (start - end) / start
```

For the case values `87.0274` and `32.2948`, the result is `62.89%`.

## Module 5 — Gates and evidence

A saved artifact is not yet a verified result. Verification should occur after
serialization in a separate process. That catches failures masked by in-memory
state and demonstrates that the stored artifact can be loaded and inspected.

Three evidence layers should be kept distinct:

1. **Integrity evidence:** source identity, artifact manifest, file digests,
   tensor inventory, shapes, data types, and finite-value checks.
2. **Intervention evidence:** targets changed as expected, non-target tensors
   remained identical, norms were restored within tolerance, and pass metrics
   were recorded.
3. **Behavioral evidence:** defined benign controls, distributional evaluation,
   capability tasks, false-refusal checks, and red-line testing.

An exact-match control establishes that the selected control was identical
under the defined comparison. It does not establish that every untested behavior
was identical.

Write limitations next to results. Do not hide the evidence boundary in a
distant appendix.

## Module 6 — Qwen 9B case file

The sanitized historical case records a three-pass DWM result on the exact base
revision named in `CASE-FILE.json`. The measurement instrument contained 32
pairs and 64 rows. The edit plan used strength `1.0` for three recaptured passes.

Recorded final-arm evidence:

- 60 of 60 target tensors changed;
- zero declared targets were missed;
- zero changed tensors were recorded outside the target manifest;
- zero non-finite target values were recorded;
- 15 of 15 defined benign controls were exact matches; and
- separation moved `87.0274 -> 50.6634 -> 41.4297 -> 32.2948`, a 62.89%
  reduction from the initial measurement.

This supports a precise conclusion: the recorded procedure changed every
declared target, recorded no change outside the manifest, preserved the defined
exact-match controls, produced finite stored values, and reduced the defined
separation metric by 62.89% over three recaptured passes.

The case does **not** include a fresh final-arm distribution study, broad
capability suite, false-refusal study, red-line evaluation, or complete release
qualification. Results from a different edit strength do not transfer
automatically.

## Capstone

Write a four-sentence decision note:

1. State the exact artifact identity and edit plan.
2. State the integrity and intervention evidence.
3. State the measured geometry result.
4. State the missing evaluation and the next decision gate.

Then complete the knowledge check in `STUDENT-WORKBOOK.md` and compare your
reasoning with `INSTRUCTOR-GUIDE.md`.
