Article content
“Modern AI infrastructure needs a balanced CPU plus GPU platform. AMD EPYC processors provide host-node performance, memory bandwidth, high-performance I/O, efficiency, and GPU ecosystem compatibility for GPU-accelerated AI systems, while AMD Instinct GPUs deliver the acceleration needed for large-scale inference. Embedded LLM’s TokenVisor Spaces show how ecosystem software can bring these two planes together for production agentic AI on AMD-powered infrastructure.”
Article content
What TokenVisor Spaces Provides
TokenVisor Spaces gives each agent a controlled workspace on customer-managed infrastructure. Each capability answers something agents need that raw infrastructure does not provide:
Article content
- Agents run for days, not single requests. Spaces provides persistent x86-64 workspaces for long-running, multi-turn agents.
- Agents execute real code against real systems. Spaces provides isolated execution for shell, files, code, browser automation, and tools.
- Agents take actions enterprises cannot let run unattended. Spaces provides human approval workflows for sensitive actions.
- Agents fail in ways someone must be able to reconstruct. Spaces provides replayable event history for review, debugging, audit, and evaluation.
- Agents consume models continuously. Spaces provides governed model access through TokenVisor, including agent guardrails, policy, metering, routing, and usage controls.
Article content
Article content
Agentic inference is making data movement part of the performance path.
Long-context agents need their accumulated context back on every turn without paying full prefill each time. The market evidence is already public: SemiAnalysis reported that across 1.5M+ of its own Claude Code requests, roughly 95% of all tokens were cache reads, cutting its prompt-token bill by about 84%. Agentic inference economics are cache economics – and the cache has to live somewhere with more capacity than GPU HBM.
Article content
Embedded LLM has validated the model-plane data path for exactly this on AMD Instinct MI355X. In a production-shaped synthetic agentic replay, storage-backed KV-cache reuse using vLLM, a KV-cache management layer, AMD hipFile/GDS, and native local NVMe delivered 3.31x lower warm-turn median latency and 2.23x faster total wall-clock time versus a matched vLLM HBM prefix-cache baseline.*
Article content
Embedded LLM is collaborating with VAST Data and Tensormesh for platform-scale KV-cache reuse, agent state, trace capture, replay/evaluation, and RL data.
Article content
“KV-cache reuse is what makes long-running, multi-turn agents economically viable,” said Kuntai Du, Chief Scientist and Co-founder at Tensormesh. “By adopting LMCache, Embedded LLM brings that infrastructure to more of the AI cloud market, giving agents persistent, reusable context so they run faster, are more cost-effective, save energy and stay auditable at scale.”
Article content
Article content
“Agentic AI requires infrastructure that can efficiently manage context, data, and state across long-running AI workflows,” said Anat Heilper, Director of AI Architecture at VAST Data. “Our collaboration with Embedded LLM helps bring the VAST AI Operating System to AMD-powered AI environments, giving customers the persistent data foundation they need to scale production AI with greater performance and efficiency.”
Article content
Availability
TokenVisor Spaces is available for partner deployment and evaluation by AI cloud operators, private AI environments, and on-premises enterprises. Prospective customers can request a demo, start a trial, join an evaluation program, or arrange a proof of concept through the Embedded LLM contact form or by emailing [email protected]. Product information is available at embeddedllm.com.
Article content
About Embedded LLM
Embedded LLM is an agentic inference infrastructure company in the AMD Instinct AI & HPC Software Ecosystem and a Red Hat ecosystem partner. The company helps AI clouds and enterprises turn GPU fleets into production AI services through vLLM-based serving, TokenVisor for governed model APIs and monetization, TokenVisor Spaces for stateful agent execution, JamAI Base for traceable AI operations, and operator-level RL/post-training infrastructure.

1 hour ago
4
English (US)