Skip to content
AI.info

Research

LLMs are General Asynchronous Agents

Overview Research area: LLM agents / inference systems — specifically asynchronous (concurrent) LLM agents, spanning streaming video understanding, voice-style interruption handling, embodied agents,

LLMs are General Asynchronous Agents
arXiv
2609.35427
Published
2026-09-28
Authors
George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko

AI summary

Overview

Research area: LLM agents / inference systems — specifically asynchronous (concurrent) LLM agents, spanning streaming video understanding, voice-style interruption handling, embodied agents, computer-use agents, and parallel reasoning.

Technical level: Advanced. The paper mixes agent design with low-level inference engineering (attention KV caches, Gated Delta Network recurrent states, rotary position embeddings, batched GPU scheduling).

Scope in one sentence: The paper proposes AsyncLLM, a training-free async/await framework and inference engine that lets modern hybrid and multimodal LLMs run multiple overlapping inference coroutines with shared memory views, and tests it on asynchronous text/image edits, streaming video, videogames, and system monitoring.

What This Paper Is About

LLM agents normally run in a turn-based loop: read, think, act, observe, repeat. Many real-world settings are not turn-based — a voice assistant must listen while thinking, a self-driving car must react to traffic changes mid-reasoning, and a monitoring agent must answer while logs keep arriving. The paper asks whether asynchrony can be treated as a general LLM capability, like tool use or few-shot learning, rather than being rebuilt as a separate specialized architecture for each domain, and answers by building a general framework in which LLMs run concurrent coroutines that share memory states.

Key Contributions

  1. AsyncLLM framework. A general, training-free framework for asynchronous LLM agents built on the async/await programming model using Python's asyncio. It lets coroutines run concurrent LLM inference while communicating through overlapping memory blocks (CacheBlocks) containing attention KV caches and Gated Delta Network (GDN) recurrent states.

  2. Parallel GPU inference with shared memory states. An algorithm for running many instances of the same LLM concurrently while each sees the others' progress in real time, supporting hybrid (full + linear attention) and multimodal LLMs. It extends prior attention-cache manipulation: full-attention layers use query rotation with RoPE, linear attention layers compose block-wise affine transitions, and multimodal layers use MRoPE correction with maintained per-block position spans.

  3. An inference engine with balanced batching. The engine (built over mini-SGLang, a minimal SGLang implementation) gathers forward-pass requests from all coroutines into a shared queue, forms balanced batches so fast reaction coroutines are not clogged by large background requests, and uses chunked prefilling plus Paged Attention.

  4. A generality demonstration across four task families. Without task-specific fine-tuning, agents built on the same Qwen 3.x model family handle asynchronous text clarifications, asynchronous image changes, streaming video understanding, videogames, and text-based system monitoring — plus a preliminary test of agents defining their own coroutines.

Main Findings

  • LLMs can act asynchronously with no fine-tuning. Qwen 3.x models can handle concurrent sub-tasks and react to new stimuli under the AsyncLLM programming model without task-specific asynchronous training.

  • Single asynchronous text inputs work. On MATH-500-Sharded (500 math problems from Hendrycks et al., each split into an incomplete prompt plus a clarification delivered after k reasoning steps), AsyncLLM agents react to mid-reasoning clarifications; hybrid AsyncReasoning baselines are reported for reference.

  • Single asynchronous image changes work. On 513 constructed image pairs from existing visual QA and math datasets, where an edited erroneous image is replaced by the original after k decoding steps, agents revise ongoing reasoning. Accuracy drops somewhat faster than for text, which the authors attribute to visual tasks needing less reasoning, so larger k sometimes arrives after the answer is already produced.

  • Training-free streaming video beats a specialist model. On SoccerNet-Caption using the streaming evaluation protocol of Ding et al., and on ProactiveVideoQA using the official protocol (PAUC at ω = 0.5), AsyncLLM with Qwen3.5+ models outperforms Mage-VL, a model trained specifically for video streaming, on both benchmarks. Mage-VL was trained with different priorities, but the larger Qwen3.5+ models also give more accurate descriptions on ProactiveVideoQA. The agent has five concurrent components (event probe, background thinking thread, output probe, description writer, output compiler) and uses probes that run the base LLM as-is with a pre-filled prompt rather than training separate probe modules.

  • Interactive videogame agents react faster than sequential agents. On two ViZDoom scenarios (HealthGathering and DeadlyCorridor, frame skip 4, averaged over 100 episodes), the AsyncLLM agent using Qwen3.6-35B-A3B reacts much faster than a sequential agent based on the same model while preserving gains from reasoning. The authors explicitly state this shows basic capability, not state-of-the-art performance for those environments.

  • System monitoring is faster at comparable accuracy. On the 34-task Monitoring subset of DevOps-Gym, with logs streamed in chunks split by lines up to 1 second or 1024 characters, AsyncLLM agents detect anomalies with comparable accuracy but far fewer forward passes than sequential baselines. For Qwen3.6-35B-A3B, the sequential baseline uses 8452 forward passes at 61.76 accuracy, whereas AsyncLLM uses 3837 forward passes at 55.88 accuracy; a skip-rows baseline drops to 26.47 accuracy. For Qwen3.5-9B, the sequential baseline uses 8726 forward passes at 47.06 accuracy versus AsyncLLM at 2582 forward passes and 41.18 accuracy, while skip-rows falls to 8.82. The monitoring agent has on average 2.06 active coroutines.

  • Throughput scales with concurrent coroutines. On 1x H200, decoding throughput in tokens per second for Qwen3.5-9B is 106 with 1 coroutine, 192 with 2, 339 with 4, and 552 with 8; for the 27B model 43, 80, 144, 239; for 35B-A3B 96, 169, 284, 438. Real-time rate on ProactiveVideoQA, in average stream seconds per GPU-second, is 4.4 / 2.9 / 2.5 for video only and 1.8 / 1.4 / 1.4 with audio (TV) for the 9B, 27B, and 35B-A3B models respectively.

  • Anti-distraction prompting is needed. Without a prompt restricting the agent to resource-usage issues and limiting its reasoning, Qwen 3.6-35B-A3B overthought unrelated issues, fell behind real-time logs, and dropped below 10% accuracy.

  • Self-defined coroutines are promising but unreliable. Letting agents define their own coroutines from the environment and the AsyncLLM API gave good initial results on HealthGathering but inferior results on DeadlyCorridor. The authors state the capability is not yet reliable.

Methodology in Plain English

The researchers treat concurrency as a programming problem rather than a model-training problem. A developer (or the agent itself) writes several independent tasks — called coroutines — that each run their own LLM inference.

Instead of exchanging text tokens with each other, the coroutines share memory. Each forward pass writes its internal state into a "cache block" holding the model's memory for a continuous chunk of tokens (attention KV caches plus GDN recurrent states). A coroutine can read several blocks at once in any order; the paper calls these compositions "cache views." If two coroutines look at each other's blocks while both are running, the resulting state does not correspond to any sequential inference, but the authors show it remains legible to the model.

To make this cheap, the engine avoids re-encoding tokens whenever a view changes. For full-attention layers it adopts prior work where only the current queries are rotated rather than all past keys and values. For linear attention layers such as Gated Delta Nets, the authors reformulate updates from per-token to per-block, summarizing each block with a pair of matrices (Â, B̂) so block composition costs operations proportional to the number of blocks rather than the number of tokens. For multimodal inputs they rotate only the temporal axis and track per-block position spans instead of token counts.

Synchronization between coroutines uses ordinary asyncio primitives — events, locks, queues — so an agent can wait for a "paragraph finished" signal, a new video frame, or a log chunk. The engine collects all pending requests, forms balanced batches (using chunked prefill so long prefills can run alongside decoding), and serves them with cache views, so a quick reaction coroutine is not blocked behind a heavy background one.

Evaluation then isolates each claim in turn: single text clarifications, single image changes, continuous video streams, interactive gameplay, and log monitoring — all using the same model family and no task-specific fine-tuning.

Why This Matters

Research impact. The paper argues that asynchrony should be treated as a general capability instead of a collection of domain-specific architectures. It offers a shared abstraction (streams of concurrent coroutines + shared memory views) and the inference algorithms needed to make it work on modern hybrid and multimodal models, which could let researchers reuse one framework where voice assistants, VLAs, and streaming video models each currently need their own design.

Real-world applications:

  • Real-time voice assistants that listen, think, and speak concurrently and handle interruptions without specialized training.
  • Streaming video monitoring, such as industrial incident detection, driving assistance, or sports commentary, where frames arrive faster than the model can process them.
  • System and operations monitoring, where an agent must react to logs of memory leaks or runaway processes while the system keeps running.
  • Interactive and computer-use agents that act on games, GUIs, or virtual environments while reasoning continues in the background.

Industry relevance. The framework is training-free, so practitioners can adapt existing open-weight models to new concurrency scenarios or combine several scenarios at once, avoiding specialized datasets and expensive fine-tuning. Efficiency figures (throughput scaling with concurrent coroutines, and stream-seconds per GPU-second) suggest batching many coroutines together can raise GPU utilization rather than fragmenting it, and the authors point to deployment possibilities such as real-time asynchronous APIs where the agent definition is supplied through the API or written by the model itself from a user prompt.

Future Directions

  • Training models to be better general asynchronous agents, in the same way current LLMs are trained to be better general tool users.
  • Improving agents' ability to write and revise their own coroutines for a given scenario, given that the current self-defined-coroutine experiments succeeded on HealthGathering but were inferior on DeadlyCorridor and are described as not yet reliable.
  • Extending the inference algorithms to more architecture variants, which the authors place in an appendix, and further characterizing latency and throughput under synthetic workloads.
  • Pushing beyond basic capability toward competitive performance, since the videogame results are explicitly framed as demonstrating capability rather than state-of-the-art results for those environments, and the monitoring agents trade some accuracy for speed.

Target Audience

Researchers and engineers working on LLM agents, real-time or streaming inference systems, and multimodal deployment; practitioners who need concurrency handling (voice, video, logs, GUIs) without collecting domain-specific data or fine-tuning; and systems engineers interested in batched inference with shared KV caches and linear-attention state composition. Readers wanting only the conceptual idea of general asynchronous agents can read the introduction and experiments, while the methodology sections require familiarity with attention caches, rotary embeddings, and GPU serving frameworks such as SGLang.

Authors’ abstract

Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

Read the original paper