Results
Learning, transfer and retention
Our experiments show where performance improves, where earlier abilities are lost and what remains unproven.
Studies reported through 11 September 2026. Results use different tasks and scoring rules, so they cannot be combined into one overall success rate. Independent replication is still needed.
Mathematics studies
Four related studies tested how training examples and answer requirements affected mathematical performance. The sequence began with calculation steps in the prompt, then tested answers without those steps and revised training prompts that had contained correction text.
Each evaluation used new numerical values within problem types represented in training. Accepted answers passed exact calculation checks and a separate AI review, except for the focused place-value study, which used calculation checks alone.
New values in trained problem types
Trained adapter
43 / 64
Base model
0 / 64
Mismatched-answer control
0 / 64
4 September 2026 · 16 trained problem types. These results show performance within the studied types.
Harder mathematics, plain-text answers
Trained adapter
2 / 81
Mismatched-answer control
2 / 81
5 September 2026 · A separate comparison. Difficulty and answer requirements changed, so the cause of the difference remains unclear.
Earlier comparisons
With calculation steps supplied, the trained adapter passed 94 of 128 questions, compared with 73 for the base model and 52 for the prediction adapter. One reply failed an additional check, so the study did not fully pass.
Without supplied steps, the adapter passed 25 of 64 questions, while both controls passed none. Repeated negative constants exceeded the study’s limit. Rebuilding the prompts led to a focused place-value result of 18 out of 18, followed by the 43-out-of-64 comparison above.
Further training and retention
A later response update produced 12 gains, 17 losses and 55 ties across 84 paired comparisons: 42 questions, each tested twice. Seven losses involved earlier skills. Neither update was adopted.
Prediction scores also worsened on new and earlier test examples. Lower training or validation loss by itself did not establish better answers or preservation of earlier abilities.
12
Gains
17
Losses
55
Unchanged
Learning and reusing rules
In a hidden-rule study, the system correctly identified 2 of 6 rules and answered 3 of 12 later questions using its saved descriptions. The full learning-and-reuse criterion was not met.
These tests kept model weights fixed and supplied the saved rules in the prompt. They concern using stored information, rather than evidence that parameter training produced lasting learning.
An unfinished comparison
The latest preparatory round collected 175 of 336 planned measurements. Two earlier rounds collected all their data but found too few questions meeting the study’s selection criteria.
The main comparison was not completed. These records therefore do not show whether the proposed approach improves learning.
What the findings support
The strongest reported mathematics result is limited to new numbers within trained problem types. The harder comparison found no advantage over its control, and later updates produced losses on earlier tasks.
Questions drawn from shared templates and repeated tests are not independent training replications. The present evidence does not establish reliable continual learning or broad transfer. These limits guide the next study.
Loopseed
A research programme at Dilate Technologies.
© 2026 Dilate Technologies