RL-0: exact policy learning from bandit reward #7

Merged
perishadmin merged 1 commit from land/emlearn-rl-zero into main 2026-08-03 18:26:40 +00:00
Contributor

Adds exact rational EML policy parameters, a three-channel Banditron with one-shot correctness feedback, deterministic exploration, held-out radial transfer, frozen-policy ablation, CLI evidence, and explicit claim limits.

Adds exact rational EML policy parameters, a three-channel Banditron with one-shot correctness feedback, deterministic exploration, held-out radial transfer, frozen-policy ablation, CLI evidence, and explicit claim limits.
RL-0: exact policy learning from bandit reward
All checks were successful
guard / guard (pull_request) Successful in 1m39s
guard / guard (push) Successful in 1m35s
4a5d721ac4
Adds exact rational EML policy parameters, a three-channel Banditron with one-shot correctness feedback, deterministic exploration, held-out radial transfer, frozen-policy ablation, CLI evidence, and explicit claim limits.

Land-Source: emlearn-rl-zero@b55afbb449c1a27e1e86e90757dfabd229f29368
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
PerishLab/emlearn!7
No description provided.