By leveraging predefined architecture templates specialized for heterogeneous dataflows, ADOR optimally balances throughput and latency. It efficiently explores design spaces to suggest architectures
that meet the requirements of both vendors and users. ADOR demonstrates substantial performance improvements, achieving 2.51× higher QoS and 4.01× better area efficiency compared to the A100 at high batch sizes, making it a robust solution for scalable and cost-effective LLM serving.
