Updating each weight inside the backward walk frees each gradient immediately, cutting peak training memory.
Replacing linear loss combinations with constrained optimization makes optimisation more interpretable.
Adding a torch C++ interface to the ALE.
Adding support for multiple ROMs in the ALE.