RL-0: exact policy learning from bandit reward #7
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "land/emlearn-rl-zero"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Adds exact rational EML policy parameters, a three-channel Banditron with one-shot correctness feedback, deterministic exploration, held-out radial transfer, frozen-policy ablation, CLI evidence, and explicit claim limits.