Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
Abstract
As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) execute multi-step reasoning, invoke diagnostic tools, retrieve domain knowledge, and coordinate across agent teams, are increasingly embedded across the edge-cloud continuum. While the biological brain accomplishes complex cognition on an exceptionally modest metabolic power budget of approximately 20W contemporary LLMs are profoundly energy- and memory-intensive, making sustainable lifecycle orchestration a critical operational priority. However, existing AI lifecycle metrics evaluate only isolated, single-model inferences or overlook multi-agent execution graphs entirely. Consequently, network operators lack foundational models to determine whether distributed agent communication incurs meaningful energy costs and where across edge-cloud tiers agent teams should physically reside. To address this gap, we introduce agentic-eCAL, generalizing the Energy Cost of AI Lifecycle (eCAL) metric to directed multi-agent workflows by coupling a closed-form two-rate single-call energy model (compute-bound prefill and memory-bound decode) with 7-layer OSI data transport. Grounded in hundreds of GPU benchmark configurations on NVIDIA A100 and H100, 16 open-weight models and 8 orchestration topologies, we validate components of the metric and study workflow placement implications. Our findings demonstrate that inter-agent text transport incurs 0.25% of workflow energy across 5G RAN, metro, and optical links. Therefore in edge-cloud agent placement the dominant energy cost of distribution is often not the transmission of inter-agent text itself, but the additional inference and context processing induced by that communication.
Community
Where should AI agent teams physically live, at the network edge or in the cloud? We extend eCAL (our lifecycle energy metric for AI, arXiv:2408.00540 (https://arxiv.org/abs/2408.00540)) to multi-agent workflows and find that agents' messages to each other cost almost nothing to send. What's expensive is the thinking those messages trigger.
What's new: agentic-eCAL
- A lifecycle energy metric for directed multi-agent workflows, not just single-model inference
- Combines a closed-form two-rate energy model for each LLM call (compute-bound prefill + memory-bound decode) with full 7-layer OSI data-transport costs
- Grounded in hundreds of GPU benchmark configurations on A100/H100, 16 open-weight models and 8 orchestration topologies
Key finding
Sending inter-agent text over 5G RAN, metro and optical links accounts for only ~0.25% of workflow energy. So when you distribute agents across the edge–cloud continuum, the real energy cost isn't the network. It's the extra inference and context processing each message causes at the receiving agent.
Why it matters: think of it as the electric-vs-diesel lifecycle comparison for agentic AI. You can't judge a deployment by one component; you have to count the whole workflow.
🌐 Project: https://www.agenticscaling.ai
Happy to answer questions about the energy model, the placement results or how to apply agentic-eCAL to your own workflows!
Get this paper in your agent:
hf papers read 2609.18283 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper