Architecture

ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput

Junsoo Kim, Hunjong Lee, Geonwoo Ko, Gyubin Choi, Seri Ham, Seongmin Hong, Joo-Young Kim

OVERVIEW

ADOR is a framework that automatically identifies and recommends hardware architectures tailored to LLM serving

By leveraging predefined architecture templates specialized for heterogeneous dataflows, ADOR optimally balances throughput and latency. It efficiently explores design spaces to suggest architectures
that meet the requirements of both vendors and users. ADOR demonstrates substantial performance improvements, achieving 2.51× higher QoS and 4.01× better area efficiency compared to the A100 at high batch sizes, making it a robust solution for scalable and cost-effective LLM serving.