RWKV
1 update on RWKV.
RWKV trains like a transformer and runs with constant memory per token
The paper releases pretrained RNN weights from 169M to 14B parameters and reports constant time and memory per token during inference.
Cutting edge on edge devices.
1 update on RWKV.
The paper releases pretrained RNN weights from 169M to 14B parameters and reports constant time and memory per token during inference.