Our proposed research

Can a system learn when to help—and when to wait?

Wavelength is preparing proof-of-mechanism research into a system that could infer why a learner is stuck, reason through uncertainty, and decide whether support—or productive struggle—is more useful.

The underlying innovation

Expert scaffolding may reflect a latent objective that can be recovered from clinician behavior.

Instead of asking experts to write every rule in advance, the proposed research uses inverse reinforcement learning to study their moment-to-moment decisions and recover the objective guiding those decisions. That scientific claim is testable—and it may be wrong.

A belief-state architecture

Perceive. Believe. Act.

The proposed loop separates what the system observes from what it infers and what it decides to do.

01

Perceive

Encode privacy-conscious interaction traces: choices, response timing, revisions, and cursor dwell.

02

Believe

Maintain a calibrated probability distribution over possible learning bottlenecks, personalized to a learner's baseline.

03

Act

Select whether and how to scaffold using uncertainty and the recovered expert objective.

Expert demonstrationsRecovered objectiveTestable policy

Falsifiable by design

Three questions decide whether the idea works.

A Phase I result could be negative. Clear failure criteria are part of the research plan.

H1

Recoverability

Does the recovered objective predict held-out expert decisions better than a rule-based baseline?

Fails if it cannot reach ≥0.70 held-out agreement.
H2

Generality

Does the learned objective transfer across clinicians and unfamiliar scenarios?

Fails if it overfits one expert or scenario family.
H3

Mechanism

Does the recovered objective produce more flexible responding than a hand-coded reward?

Fails if simpler authored rewards perform as well.

Beyond standard adaptive branching

What makes the proposed work technically different.

Infer why—not only what

Similar mistakes can have different causes. The system must separate uncertainty, comprehension, attention, and strategy.

Beliefs, not fixed labels

The architecture represents uncertainty and updates it as new evidence arrives instead of declaring one permanent state.

Recovered, not authored

The scaffolding objective is learned from expert decisions rather than reduced to a predetermined decision tree.

Preserve learner agency

The goal is not maximum intervention. Sometimes the best action is to wait and let productive struggle continue.

High-risk technical work

The difficult parts are measurable.

01

Reward recovery is underdetermined.

Multiple objectives can explain the same expert action. We must test held-out agreement and transfer across clinicians.

02

Latent learner states may not be separable.

The target is bottleneck classification F1 ≥0.70 with expert inter-rater agreement κ ≥0.80.

03

Signals can be ambiguous.

Response delay, revision, and dwell time can have several causes, so causal identification must survive held-out perturbations.

Current stage: pre–Phase I groundwork. Wavelength has defined the population, scientific framework, conceptual workflow, clinical partnership, and patent landscape. No inference engine or decision model has been built or validated, and no clinical efficacy claim is made. The proposed system is educational technology—not a diagnostic tool or a substitute for individualized clinical care.

Build with us

Help shape what social learning can become.

We welcome conversations with families, educators, clinicians, researchers, and development partners.

Start a conversation