NVIDIA used IFA Berlin this week to push local AI further into ordinary households, and the most interesting part of the announcement isn't a new GPU. It's a free tool called NVIDIA PAIR that turns every RTX PC already sitting in a house into shared compute for AI agents.

The pitch behind Personal AI Router is simple: more than half of US households already own two or more PCs, and most of that hardware sits idle for most of the day. PAIR automatically finds compatible RTX machines on the same local network and routes AI inference jobs to whichever one has spare capacity, instead of making a single GPU grind through every request in sequence. It works with Ollama and LM Studio, adjusts as devices join or leave the network, and NVIDIA gave a concrete example: an agent sorting a cluttered inbox can split that work into subagents and have PAIR spread the jobs across multiple machines at once rather than queuing them on one.

PAIR runs in beta on Windows, macOS and Linux, and supports GeForce RTX 20 Series cards and newer, RTX PRO workstation GPUs from Turing onward, DGX Spark, and Apple M4 silicon. That cross-platform reach, including Apple hardware, is notable for a tool NVIDIA is giving away for free.

The Bigger Shift: Local AI Getting Rid of Its Setup Tax

Running an AI agent locally instead of through a cloud subscription has always meant real friction: picking a model, finding a compatible inference server, tuning quantization settings, keeping everything updated. NVIDIA says three widely used agent apps are removing most of that friction on Windows. Hermes Agent, built by Nous Research, gets one-click setup that auto-detects the GPU and configures itself through llama.cpp. OpenClaw, the largest AI project on GitHub with more than 380,000 stars, is getting a Windows app built with Microsoft to simplify local setup on any RTX GPU with at least 24GB of VRAM.

The more telling move is Perplexity's. The company built its business on cloud-hosted answers, and it's now shipping a Portable Computer agent that runs complete workflows locally on RTX GPUs with at least 24GB of VRAM, without spending credits, and only escalates to a cloud model when a task genuinely needs one. An AI company that sells cloud access is building a product whose selling point is not needing the cloud, which says something about where user demand for privacy and cost control is actually pointing.

NVIDIA is also claiming real inference gains, not just easier setup: up to 1.9x higher throughput on llama.cpp from kernel optimizations on an RTX 5090, and up to 1.4x on two clustered DGX Spark units running vLLM. Both are available now through LM Studio and Ollama, not locked behind a future driver release.

New Hardware Arrives in October

The IFA news also confirmed new NVIDIA RTX Spark systems. Acer showed a compact desktop concept, and Lenovo announced the Yoga Pro 9n laptop and Yoga 9n 2-in-1, joining the six OEMs already committed to shipping RTX Spark hardware in October. RTX Spark pairs a 1 petaflop RTX Blackwell GPU with up to 128GB of unified memory and a 20-core Grace CPU, aimed at running always-on agents locally rather than just accelerating games or creative apps.

CyberLink is also bringing generative editing tools to PhotoDirector 365 through a new AI PC Mode built for RTX Spark, covering background removal, object removal and portrait refinement with the option to keep processing on-device instead of sending images to the cloud.

None of this is available in the Philippines yet in a form buyers can walk into a store and pick up, but it signals where NVIDIA expects the next wave of AI PC purchases to matter: not raw benchmark numbers, but whether the hardware someone already owns can quietly do more.