Nathan Wilce

sequence models · inference

I spend my time optimizing the algorithms powering the multi-model inference system we run at Featherless.ai: the cluster itself, every layer of hardware and software, and the load-distribution that decides which model sits where. I also spend a lot of time trying new ideas in LLM architecture, inference and more.

S ← wS + v⊗k