Neumeral Labs
Big models are easy. Small, fast, smart ones are the hard part.
Neumeral Labs is our research group for practical machine learning — getting real intelligence to run on constrained hardware, cutting inference latency, and chaining models into systems that outperform any single one of them.
Overview
Research aimed at production, not a leaderboard
Most published AI research chases benchmark scores on hardware nobody ships to customers. Neumeral Labs starts from the constraints our clients actually have: limited RAM, tight latency budgets, no GPU, no reliable connection.
We treat research as applied work. Anything that survives contact with a real device or a real cost budget gets folded back into what Neumeral builds for clients — and anything that doesn't, we write up so nobody else has to rediscover it the hard way.
See what we're exploringProblems we're interested in
What we're exploring
Fit real models on hardware that barely has room for them
Most inference stacks assume gigabytes of RAM and a GPU nearby. We're studying quantization, pruning and memory-mapped weight streaming to get useful models running on microcontrollers, budget phones and other devices with a few hundred megabytes to spare.
Talk to us about thisShave milliseconds out of every forward pass
Speculative decoding, KV-cache tricks, batching strategies and kernel-level optimizations each buy back latency. We benchmark these techniques against real workloads to find which ones actually move the needle, and which are just leaderboard tricks.
Talk to us about thisChain small models into something smarter than one big one
A router that picks the right specialist, a cascade that escalates only when needed, an ensemble that cross-checks itself — compound systems can beat a single large model at a fraction of the cost. We're mapping which architectures earn their complexity.
Talk to us about thisKeep the model on the machine that's asking the question
Not every workload should phone home to an API. We're building and benchmarking local-inference setups — on laptops, workstations and on-prem servers — for teams that need privacy, offline reliability or predictable cost instead of a per-token bill.
Talk to us about thisHow we work
Explore, benchmark, then ship what holds up
Explore
We track techniques worth testing — from academic papers to obscure GitHub repos — and pick the ones with a real shot at production use.
Benchmark
We measure against the constraints that matter — RAM footprint, latency under load, accuracy trade-offs — not just a leaderboard number.
Apply or publish
What holds up gets folded into client engagements. What we learn along the way, we write up so the next team doesn't have to re-run the experiment.
It's the one that runs where you need it.
Interested in the research?
We're early, and looking for problems worth solving alongside partners who feel these constraints for real — tight hardware, tight latency, tight budgets. If that's you, get in touch.
Get in touch