sequence models · inference
I spend my time optimizing the algorithms powering the multi-model inference system we run at Featherless.ai: the cluster itself, every layer of hardware and software, and the load-distribution that decides which model sits where. I also spend a lot of time trying new ideas in LLM architecture, inference and more.
2025 · arXiv:2503.14456 · generalized delta rule
multi-model inference · load distribution · infra