BEFORE YOU BEGIN
Read this as an evidence course.
DWM is the method under study, but the transferable skill is disciplined inference. You will repeatedly ask: What was fixed? What changed? What was measured? Which controls held? What conclusion is supported—and where does the evidence stop?
This material is educational. It contains no model weights, raw paired bank, activation tensors, or execution code. Work only with systems and artifacts you own or are explicitly authorized to evaluate.
SCOPE AND VOCABULARY
Name the intervention before judging it.
Domain Weight Modification is a targeted model-weight intervention guided by measured activation differences. Its purpose is to modify a defined geometric relationship in selected weights while keeping the procedure bounded and auditable.
Four methods that should not be collapsed
Measures a direction from paired observations and applies a constrained update to selected matrices.
Optimizes model parameters against a training objective, typically using gradient-based learning over a dataset.
Removes, disables, or zeros a component to test its causal role. That is not the update studied here.
Changes runtime context or inputs without changing the stored model weights.
Objective, method, evidence, conclusion
Keep these four layers separate. The objective says what relationship you intend to change. The method says exactly how. The evidence records what the instruments observed. The conclusion must remain inside the evidence boundary.
If a geometric separation metric decreases, which statement is supportable?
- The defined metric decreased under the recorded measurement procedure.
- The model is universally safer.
- No capability changed anywhere.
Answer: 1. The other claims require independent evidence.
MEASUREMENT INSTRUMENT
Freeze the identity before measuring change.
A result cannot be reproduced from a model name alone. The experiment identity is a bundle: exact base revision, tokenizer behavior, paired bank, prompt construction, capture point, target matrices, numerical precision, algorithm parameters, and evaluation rules.
The paired bank is an instrument
A paired bank presents two defined variants for each concept under comparison. In a prefill-only procedure, activations are captured from the supplied token sequence without sampling a continuation. The difference between paired representations contributes to a direction estimate.
Once accepted, the bank should be immutable for that experiment. Record a cryptographic digest and reject any silent change. If a pair is rewritten halfway through the process, the measuring instrument changed too.
Minimum identity record
- Exact model and source revision
- Tokenizer and chat-template behavior
- Bank row count, pair count, and SHA-256 digest
- Capture point and token-position rule
- Selected layers and matrix families
- Pass count, strength parameter, precision, and norm policy
- Gate thresholds and independent verification procedure
If the revision, digest, architecture map, or parameters do not match the approved identity, stop. Do not “continue carefully” under a changed instrument.
ARCHITECTURE AND EDIT SURFACE
Map the surface you intend to change.
The direction lives in residual-stream space. An edit affects a matrix only when that matrix consumes a compatible representation, so architecture inspection is part of the method—not a setup detail.
Conceptual data path
For a matrix W with columns in the residual dimension and a unit direction d, a rank-1 projection can remove the component of each row aligned with d. The exact orientation must be verified against the implementation’s weight layout.
Here, α is the edit strength. A target manifest should enumerate the exact tensors, expected shapes, and expected count. A mismatch is a gate failure, not an invitation to guess.
Architecture gate questions
- 1
Does the captured representation dimension match the edited matrix input dimension?
- 2
Are all expected target matrices present exactly once?
- 3
Are excluded matrices demonstrably unchanged?
- 4
Does serialization preserve the intended tensor names, shapes, and data types?
ITERATIVE RANK-1 EDITING
Measure, update, restore, then measure again.
The runbook pattern is iterative. Each pass captures fresh activations from the model produced by the preceding pass. This recapture matters because the representation geometry may move after an edit.
Capture
Run the frozen paired bank in prefill mode and collect representations at the declared capture point.
Estimate
Aggregate normalized pair differences into a unit direction for the current pass.
Update
Apply the bounded rank-1 projection to every tensor in the target manifest.
Restore
Rescale each edited matrix to its own pre-edit Frobenius norm.
Gate
Check target count, non-target identity, finite values, controls, and recorded metrics.
Recapture
If another pass is approved, estimate the next direction from the newly edited model.
Norm restoration constrains scale drift; it does not prove that all behavior is preserved. That distinction is why the workflow also needs benign controls and broader evaluation.
GEOMETRY CALCULATOR
Compute relative reduction
Use the historical case values or enter your own positive measurements.
GATES AND EVIDENCE
A saved artifact is not yet a verified result.
Verification should occur after serialization in a separate process. That catches failures hidden by in-memory state and proves the stored artifact can be loaded and inspected independently.
Three evidence layers
01Integrity evidence
Exact source revision, artifact manifest, file hashes, tensor inventory, shape checks, data types, and finite-value checks.
02Intervention evidence
Target matrices changed as expected, non-target tensors remained identical, matrix norms were restored within tolerance, and pass metrics were recorded.
03Behavioral evidence
Defined benign controls, distributional evaluation, capability tasks, false-refusal checks, and red-line testing. Each result applies only to its actual test set and conditions.
Claim ladder
“All 60 target tensors changed; the defined 15 benign controls were exact matches.”
“No non-target behavior changed.” The listed controls are evidence, but they do not cover every behavior.
“The artifact is ready for unrestricted deployment.” Release qualification requires its own evaluation and approval record.
Write the limitation next to the result, not in a distant appendix. A reader should be able to see the evidence boundary before repeating the claim.
QWEN 9B CASE FILE
Audit the record. Then write the conclusion.
This sanitized historical case records a three-pass DWM result on a pinned Qwen 9B revision. Treat it as evidence to inspect, not as a release recommendation or a substitute for fresh evaluation.
Recorded final-arm evidence
All tensors in the target manifest changed.
No declared target was missed.
No changed tensor was recorded outside the manifest.
The independent scan recorded only finite values.
Exact-match controls were identical under the defined comparison.
87.0274 → 50.6634 → 41.4297 → 32.2948.
What the case supports
The recorded procedure changed every declared target, did not record changes outside that manifest, preserved the defined exact-match controls, produced finite stored values, and reduced the defined separation metric by 62.89% over three recaptured passes.
What the case does not establish
The record does not include a fresh final-arm distribution study, broad capability suite, false-refusal study, red-line evaluation, or complete release qualification. Earlier results from a different edit strength do not transfer automatically to this arm.
CAPSTONE / WRITE THE DECISION NOTE
Use four sentences.
- State the exact artifact identity and edit plan.
- State the integrity and intervention evidence.
- State the measured geometry result.
- State the missing evaluation and the next decision gate.
A strong note is compact because its nouns are precise. Compare yours with the answer guidance in the downloadable instructor guide.
KNOWLEDGE CHECK
Test the evidence chain.
Choose one answer for each question. Your selections are not transmitted or stored.
COURSE COMPLETE
Keep the evidence boundary visible.
You now have a framework for reading a DWM record without confusing a measured geometric result with a broader behavioral or deployment claim.