Systems research · 2026
Cache Cliff edge-AI research
A hardware-aware way to choose vision encoders for edge visual-language systems.
Visit live projectSituation
Edge AI selection often prioritizes accuracy while overlooking the actual latency bottleneck.
Constraint
Recommendations needed to reflect how architectures behave on constrained hardware, not just in server benchmarks.
Decisions
Synthesized public benchmark and systems evidence into a hardware-aware selection framework for edge visual-language systems.
Outcome
Introduced the Cache Cliff terminology to make non-linear hardware trade-offs explicit; readers should validate recommendations on their target hardware.
Related Technical Research & Systems Deep-Dive
The Cache Cliff: a hardware-aware edge AI framing →
Read the empirical benchmark across 8 vision encoders and 4 hardware platforms detailing roofline limits and time-to-first-token optimization.