Back to selected work

Systems research · 2026

Cache Cliff edge-AI research

A hardware-aware way to choose vision encoders for edge visual-language systems.

Research FellowBenchmarkingEdge inferenceLatency analysis
Visit live project

Situation

Edge AI selection often prioritizes accuracy while overlooking the actual latency bottleneck.

Constraint

Recommendations needed to reflect how architectures behave on constrained hardware, not just in server benchmarks.

Decisions

Benchmarked eight architectures on four hardware platforms and isolated the encoder contribution to time-to-first-token.

Outcome

Vision encoders accounted for 70–85% of time-to-first-token; the Cache Cliff framework made hardware-aware trade-offs explicit.