JIPCET 2026; 3(2)
01 Jul 2026
pp. 73–104
Latency-Aware Cost Models for Serverless Inference at the Network Edge: Derivation and Multi-Platform Validation
Michelle Tan, Joseph Haddad, Peter Osei, Aisha Rahman
We present a latency-aware cost model for edge inference, showing modest batch sizes trade 8% throughput for a 22% reduction in per-token expense.
Read article →