Posts

Before You Spend $15,000 on Local AI

A serious local-AI budget gives you options. Memory math, quantization, and real workload testing should decide which machine you buy.

The Guy in the Hat, matching the Hat in the Loop mascot with patterned red cap and red goatee, considers local AI computers with a checklist. Headline: Start with the job.

Give a computer nerd a $15,000 budget for local AI and watch how quickly the conversation turns into a shopping cart.

Big GPU. More memory. Bigger model.

I get it. That sounds like a fun afternoon.

But before spending that money, I want to know what the machine actually needs to do.

Read service manuals? Extract information from invoices? Help write code? Support one person, or several people working at the same time?

Those answers should drive the purchase.

The memory math matters

A 96GB graphics card is a serious piece of equipment. It also has limits.

A model with 70 billion parameters needs roughly 140GB just for its weights at 16-bit precision. That is 70 billion multiplied by two bytes. It does not include the extra memory needed to run it.

So no, a 70B model at FP16 does not fit entirely on a 96GB card.

Quantization changes that equation by storing weights with fewer bits. At a theoretical four bits per parameter, those same 70B weights are about 35GB before quantization metadata and runtime overhead.

That makes larger models much more practical. It also means testing whether the compressed version still does your job well.

And leave room for the conversation. The KV cache, which stores information used during generation, also consumes memory. Longer inputs and more simultaneous requests can change what fits.

“Loads successfully” is the beginning of a test.

Three directions I would investigate

A single large-memory NVIDIA workstation.

The RTX PRO 6000 Blackwell Workstation Edition has 96GB of GDDR7 memory and 1,792GB/s of memory bandwidth. That makes it a compelling candidate for a personal local-AI workstation using quantized large models.

One GPU can simplify deployment. It does not make every model or software stack automatically compatible.

NVIDIA also lists a maximum power consumption of 600W for this card. The complete machine needs appropriate power and cooling, and the complete quote needs to fit the budget. I would get that quote before promising a $10,000–$14,000 build.

Source: NVIDIA specifications

A refurbished system with a datacenter GPU.

Worth investigating when the seller can demonstrate the exact workload and support the complete system.

Meta documents Llama 4 Scout as having 109B total parameters, with 17B active, and describes single-H100 inference using INT4 quantization. The active parameter count is not the amount of weight storage you should budget for.

That example does not establish identical performance on an A100, guarantee a particular context length, or promise a responsive multi-user system.

I would want a validated configuration, a warranty, and a workload demonstration before buying. A $15,000–$25,000 quote is mostly above a $15,000 ceiling.

Source: Meta’s Scout model card

A Mac Studio with substantial unified memory.

Also worth evaluating for a personal office setup. Apple’s unified-memory approach gives the CPU and GPU access to a shared pool, so allow room for the operating system and applications too.

The decision comes down to the exact configuration, model, runtime, and response time you need. Check current availability and pricing, then benchmark. A memory-capacity comparison alone will not tell you which machine feels faster.

Source: Apple specifications

Smaller models deserve a fair test

I would not call a 7B or 13B model a toy.

If a smaller model handles a defined task accurately, quickly, and at a lower operating cost, that is a useful result. Bigger should earn its place through better performance on the work.

Before buying, I would test representative documents and questions, measure accuracy, check time to first response, and run the number of simultaneous requests I actually expect.

For the on-prem conversations I have at Subatomic, that distinction matters. Buying hardware is one part of building something people can depend on. Data access, permissions, human review, and maintenance still need an owner.

A $15,000 budget gives you options.

Start with the job. Test the model. Then buy the machine.

Otherwise, you might end up with a very expensive space heater that writes surprisingly good emails.