Back to selected work
Systems research · 2026
Cache Cliff edge-AI research
A hardware-aware way to choose vision encoders for edge visual-language systems.
Visit live projectSituation
Edge AI selection often prioritizes accuracy while overlooking the actual latency bottleneck.
Constraint
Recommendations needed to reflect how architectures behave on constrained hardware, not just in server benchmarks.
Decisions
Benchmarked eight architectures on four hardware platforms and isolated the encoder contribution to time-to-first-token.
Outcome
Vision encoders accounted for 70–85% of time-to-first-token; the Cache Cliff framework made hardware-aware trade-offs explicit.