the techy bits

The AI doesn't live
in the cloud.

We build and run the hardware, not slide decks. There is an NVIDIA DGX Spark sitting on a desk in our shop, running the models that answer our phones, route our trucks and read our receipts — on our own power bill, behind our own front door. We own the machine, so you are not renting someone else's API by the token.

This page is the running list of what's actually on it. Everything here is live, not planned.

NVIDIA Inception Program member badge

A member of the NVIDIA Inception program

the box

One Blackwell GPU, in a shop

A DGX Spark — NVIDIA's desktop-class Blackwell machine. It is not a rack in a datacenter and it is not a virtual instance. It is a physical box, in a working service business, plugged into the wall.

Everything below runs on that one GPU, quantized to fit and served locally over an encrypted private network. There is no per-token bill attached to any of it.

the stack

What's actually loaded

Nemotron 3 33B
NVIDIA's reasoning model, 4-bit quantized. The main brain — analysis, drafting, judgment calls.~3s warm response, ~6.7s cold start.
Qwen3 32B
Second reasoning model, 4-bit quantized. A different opinion to check the first one against.
Llama 3.2 Vision 11B
Multimodal. Reads receipts and job photos and turns them into structured data.~7.2s per image on real, crumpled, badly-lit receipts.
Llama Guard 3 8B
Safety and content moderation, run locally before anything is sent to a customer.
Phi-4 Mini
Small and fast. Classification and routing — deciding which of the big models a job should go to.
Nomic Embed Text
768-dimension text embeddings for semantic search across years of operational history.
SigLIP image embeddings
Visual search over the job-photo archive. Find every photo of a specific item without anyone having tagged it.
NVIDIA cuOpt
GPU route optimization. Solves vehicle-routing-with-time-windows problems against real dispatch history.
DiffusionGemma 26B
Block-diffusion rather than autoregressive — 256 tokens generated in parallel per step. Used as a fast triage layer.~593 tok/s in-step, full response in 7–8s.
why it matters

Owning the machine, not the meter

Your data stays put. Customer names, addresses, phone numbers, payroll, financials — none of it is shipped to a third-party model provider, because there isn't one. For a lot of owners that is the difference between using AI and not being allowed to.

The cost is fixed. Metered AI punishes you exactly when it starts working — the more useful it gets, the bigger the invoice. Hardware you own costs the same whether it answers ten questions a day or ten thousand.

It's fast enough to sit inside real work. A three-second answer can live inside a dispatch decision or a customer text. A thirty-second answer can't.

the honest part

What this doesn't mean

Local hardware isn't automatically better. Frontier cloud models are still stronger at the hardest reasoning, and we use them where that's the right call. The point isn't purity — it's having both, and knowing which work belongs where.

The unglamorous truth is that most of the value in an AI system isn't the model at all. It's the plumbing: clean data, real integrations, and someone who understands the business well enough to know what's worth automating.

credentials

NVIDIA Inception

We were accepted into the NVIDIA Inception program — not for a pitch deck, but for AI that already runs a real service business every day.

Curious what this would look like in your business?

Fifteen minutes, no deck. We'll tell you where AI actually fits and where it doesn't.

Book a Call