Skip to content
XINKby eYs3D Microelectronics

Technology03

Edge inference

reasoning that lives on the device

Cloud round-trips are a luxury a moving robot doesn't have. XINK CORE runs LLM and VLM reasoning where the data is captured — private, low-latency, and still working when the network isn't.

4.6 TOPS
NPU on the eCV SoC
1B
params — LLM running in live demos
28 nm
next-gen LLM IC in tape-out

01

Why the edge wins

Latency: decisions in milliseconds, not round-trips. Privacy: video and voice never leave the device. Resilience: inference keeps running offline. For machines that act in the physical world, edge inference isn't an optimization — it's a requirement.

Device inference vs cloud round-trip

02

Edge-cloud hybrid, by design

Small models reason on-device — a quantized Gemma 3 1B handles intent parsing, privacy filtering, and instant interaction — while the cloud orchestrates heavy lifting: planning, external APIs, large-model support. Secure transport, two-way collaboration, each side doing what it's best at.

Edge-cloud hybrid architecture

03

Language meets vision

Give an LLM agent a camera — captioning, visual question answering, phrase grounding, detection — and a robot can learn its surroundings and adapt to objects it has never seen. That's what turns automation into autonomy: environments no longer need to be predictable.

VLM grounding: objects tagged in a scene

04

Next: a processor built for language

The upcoming 28 nm LLM IC targets low-context inference with efficient prefill and decode — real-time intelligence and context awareness at edge power. It completes the Sense-and-React architecture: perception and cognition as a single edge reflex.

Related

Start with a dev kit, scale to a fleet.

Get XINK SENSE on your bench this week, or talk to us about the full stack.