Agent Box mini PC targets local 100B-parameter AI models
techradar.com

Agent Box mini PC targets local 100B-parameter AI models

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAcrab has unveiled Agent Box, a compact system intended to run large language models locally, including claimed 100B-parameter models. Its value for builders is the combination of on-device inference and an agent software stack, though the headline performance claims still need independent testing.

Acrab, a Singapore-based startup, has unveiled the Agent Box mini PC, a compact system designed to run large language models locally rather than sending every request to a cloud API. The practical pitch for AI builders is lower per-token cost, tighter data control, and agents that can keep working when connectivity is limited. The important caveat is that the device's biggest claims are based on company disclosures and internal tests, not independent benchmarks. Acrab's Agent Box is designed to run 100B-parameter language models on local hardware.

A specialized SoC does the heavy lifting

Agent Box uses Acrab's proprietary G≡LIX 1 SoC, manufactured on a 5nm process. The chip combines a 20-core Arm CPU, a 3 TFLOPS GPU, and a multicore neural processing unit aimed at large language model inference. Acrab reports up to 700 TOPS, 273GB/s of memory bandwidth, an 8MB L1 cache, and a 768GB/s L2 cache.

Those specifications suggest a design optimized for moving model data and processing inference efficiently, rather than a general purpose mini PC with an AI accelerator added afterward. Acrab also says the platform includes AI runtimes, developer toolchains, operating system capabilities, reference designs, and orchestration software for autonomous agents. That software layer may matter as much as the silicon for teams trying to deploy agents outside a server environment.

The performance claim is promising, but narrow

In an internal test, Agent Box reached a prefill rate of 1,416.8 tokens per second using a Gemma 26B A4B configuration with a 40K KV cache and 10K token input. The reported result was 188.9 tokens per second on an Apple Mac Mini M4 Pro, or roughly 7.5 times faster in that test. The reported prefill comparison used a specific Gemma 26B A4B configuration and cache setup.

Builders should avoid reading this as a general 7.5x speed advantage. Prefill measures processing an existing input context, while generation speed, model quantization, context length, power limits, compiler maturity, and tool calls can change the experience of an agent. The source also does not provide an independent DGX Spark benchmark, despite Acrab positioning the system near Nvidia's $5,000 machine at about one-fifth the cost and half the power consumption.

Where local inference could change a product

A local system is useful when an application handles private data, needs predictable operating costs, or must respond without a dependable internet connection. A founder building an industrial assistant, for example, could keep sensor data and operational history on site. A developer building a household agent could use local inference for device control while limiting cloud exposure.

Acrab's demonstrations included voice commands that created and printed 3D models, controlled a robotic vacuum, and interacted with smart lights, air conditioners, and locks. These demos point to an agent harness connected to real tools, but they do not establish reliability, safety, latency under sustained workloads, or how much human approval was required.

The company also describes future applications in AI NAS products, network-attached storage, industrial and service robots, and smart vehicles through hardware partnerships. For product teams, that is an ecosystem proposal rather than an available integration roadmap. Adoption will

Sources

Latest Tech News