AA-AgentPerf-Local benchmarks inference speed on local AI models

Summary

Artificial Analysis has announced the launch of AA-AgentPerf-Local, an open-source inference testing tool designed for local AI models, allowing users to evaluate the performance of agentic AI on their personal laptops and workstations. This tool replays real agent trajectories to benchmark inference speed across several hardware platforms, including the NVIDIA DGX Spark and GeForce RTX 5090. As developers increasingly seek local evaluation tools for privacy and latency considerations, AA-AgentPerf-Local will help users make informed decisions about their AI serving configurations. The tool's initial results spotlight significant variations in completion times across different models and systems, emphasizing the growing demand for capabilities that cater to both individual and organizational AI needs.

Analysis

Qwen3.5-9B: Qwen3.5-9B is a dense open-weights language model released as part of the Qwen series. It was one of the four models featured in the initial AA-AgentPerf-Local benchmarks at 4-bit quantization to reflect realistic local serving conditions. The model serves as a baseline in comparisons of inference performance across hardware. Qwen3.8-27B: Qwen3.8-27B is a dense open-weights model from the Qwen family used for high-performance inference testing. It was benchmarked in the AA-AgentPerf-Local launch results to evaluate completion times and the impact of speculative decoding. The model helps illustrate how different architectures perform under agentic workloads on local systems. Ling 3.0 Flash: Ling 3.0 Flash is a large mixture-of-experts open-weights model with 124 billion total parameters but only a small active parameter count. It was one of the initial models tested by AA-AgentPerf-Local to cover varying memory requirements and architectures at 4-bit quantization. The model contributes to understanding performance trade-offs in local serving configurations. Qwen3.6-35B-A3B: Qwen3.6-35B-A3B is a mixture-of-experts open-weights model with a small number of active parameters relative to its total size. It was included in the AA-AgentPerf-Local initial model set and demonstrated strong speed results due to its active parameter count. The model is used to compare dense versus MoE architectures in local agent inference benchmarks. NVIDIA DGX Spark: NVIDIA DGX Spark is a compact AI system featuring unified memory and support for CUDA-based inference workloads. It was included in the initial benchmark results of AA-AgentPerf-Local as one of the tested platforms representing high-performance desktop-scale hardware. The system is positioned alongside other unified-memory options for comparing local agent inference speeds. AMD Ryzen AI Halo: AMD Ryzen AI Halo refers to the AMD Ryzen AI Max+ 395-powered system with unified memory architecture supporting ROCm for AI workloads. It was benchmarked in the launch results of AA-AgentPerf-Local alongside similar-spec systems to highlight comparative performance on agent trajectories. The platform is evaluated for its balance of memory capacity and bandwidth in local AI serving. AA-AgentPerf-Local: AA-AgentPerf-Local is an open-source inference testing tool developed by Artificial Analysis for benchmarking local AI model performance on laptops and workstations. It replays recorded real agent trajectories to measure inference speed in a standardized way across different hardware and serving configurations. The tool was announced and released to help individuals and companies evaluate setups for running agentic AI locally. MacBook Pro M5 Pro: MacBook Pro M5 Pro is an Apple laptop equipped with the M5 Pro chip, unified memory, and Metal graphics support for on-device AI tasks. It was tested in the AA-AgentPerf-Local initial hardware set as the sole laptop example, showing competitive results on certain models. The system represents portable options for running agentic AI inference. NVIDIA GeForce RTX 5090: NVIDIA GeForce RTX 5090 is a consumer graphics card with high memory bandwidth designed for demanding GPU-accelerated tasks including AI inference. It featured prominently in the AA-AgentPerf-Local initial results as the fastest system among those tested for models that fit in its VRAM. The card supports CUDA and is used to demonstrate performance advantages in single-user decoding scenarios. Agent Workloads: Agentic AI applications typically involve multi-turn interactions with growing context that stress both prefill and decode phases of inference. Hardware Platforms: Local AI serving now spans diverse compute backends including CUDA, ROCm, Vulkan, and Metal, each with distinct performance characteristics for agent tasks. Local AI Expansion: Developers and organizations are increasingly seeking tools to evaluate AI inference performance on personal hardware for privacy, latency, and cost reasons.

Categories

aimachine_learningai_agentstechvirtuals
View Original Tweet