The price of privacy: local models vs API, in numbers
A 24 GB card, a rental GPU, or an API: the invoice is not the decision. Data that cannot leave the room is.
A local assistant can parse inquiries, answer from a price list, and refuse when the file is silent. That is a measured fact, not a slogan. The remaining question is the invoice — and whether the invoice is the right question.
The choice is own silicon, a rented GPU, or a hosted API. The numbers do not all point the same way.
The floor is 24 GB
Below 24 GB of video memory there is no conversation for a 17 GB weight file plus a working context. CPU, 64 GB of RAM, disk, and a serious power supply complete the box. A new workstation in that class is a four-figure euro outlay in Western retail — typically the high four figures once a current 24 GB NVIDIA card is in the build.
AMD cards can look cheaper per gigabyte. Almost all of the inference stack assumes NVIDIA CUDA. Treat AMD as a research project, not a default for a desk that has to stay up.
The reference card for a small office is NVIDIA: a current 4090-class board, or a still-common 3090 with 24 GB.
Used silicon: about half, and what you pay instead
A used 3090 in a working workstation often lands around half of a new 4090-class build — still four figures, not a round of coffee. The discount is not free.
3090 memory is on both sides of the board. The back side is poorly cooled. Mining-era cards sat at 95–105 °C for years. Pads dry out. Solder fatigues. Dead chips get transplanted. A sealed box at a meetup proves nothing.
If the card is bought in person:
- Open it. Broken seals, oil, burnt smell, flux — walk away.
- FurMark 15–20 minutes. No artifacts, no driver death.
- GPU-Z under load: core under about 75 °C. Memory must not run past the high 80s / low 90s. Past 100 °C, cooling is dead.
- A VRAM test at ~90% fill, fifteen minutes, zero errors.
- A power-spike 3D bench so the VRMs are actually asked to work.
Rental: the market that does not quote
If you will not buy, you rent. Honest public prices for a single 24 GB guest are thinner than the marketing suggests. Large clouds sell datacenter GPUs in pairs with a high floor. Many hosts show a calculator and no total, or a “leave a request” form.
What a Western operator can actually book, in round figures:
- Hourly cloud (a 3090/4090-class guest on a spot-style host): tens of euro-cents per hour. Left on all month, that becomes low-to-mid hundreds of euro per month.
- Dedicated box with a 4090-class card: often lower per month than 24/7 hourly, because you are not paying idle mark-up to a burst market.
- Hourly only wins if the machine is actually off nights and weekends. Office hours can land around a quarter of the 24/7 hourly bill. “We will remember to shut it off” is a process, not a line item.
Power and the costs nobody puts in the spreadsheet
A box that sits warm and answers all month draws a small electricity bill — tens of euro, not hundreds, in a typical EU tariff. That is the socket.
The spreadsheet usually omits: the week to stand it up, the person who owns the driver, a dead fan, a blown PSU. Rent folds some of that into the price. Own silicon does not.
When own hardware catches the rental
Payback is a function of hours, not of faith.
- 24/7 duty. A used 24 GB box can catch a 24/7 hourly cloud bill in a small number of months. Against a cheaper dedicated rental, expect closer to three quarters. A new build stretches that toward a year or more.
- Nine-to-five. The used box may take well over a year. A new card may not catch the rental inside the useful life of the board.
These are order-of-magnitude readings on public rental menus, not a promise that a particular quote will match.
The third path, which the hardware story tries to skip
If the job is “an assistant,” the cheap path is usually not to buy or rent a GPU. It is a hosted API, billed per token.
A thousand short office turns a month on a current small or mid model is typically single-digit to low tens of euro, not hundreds. A stronger model still sits far below a GPU rental, let alone a purchase.
If the answer can wait until morning — classify last night’s inbox, file a report — batch APIs are cheaper again. The same split as in the capability bench (live chat versus bulk) divides the bill.
Own hardware is not bought to win on cost. It is bought so the data never leaves the room.
Cheap or ours
An API call sends client threads, invoices, and whatever else was in the prompt to someone else’s cluster. A local model keeps them behind the office door.
Buy the box when the load is dense and always-on, or when the documents cannot go out on any terms. That is not “expensive versus cheap.” It is cheap versus ours.
The studio will say which of the three paths fits a named volume. There is no public price list for that conversation. Start with a written brief: residency constraint, hours of duty, and whether the job is a live desk or a night batch.
- Is the real requirement data residency, or a lower token bill?
- Is the box on 24/7, or office hours — the payback story changes with the clock?
- If a used 24 GB card is in play, was it load-tested, not just unboxed?
- Is batch work on a cheaper, delayed API path, and live chat on a latency path?