On-Premise LLM for a Clinic: The Real Bill of Materials (GPU, Power Draw, and the Maintenance Nobody Quotes You)

Every vendor quoting you an on-premise LLM for your DHA-licensed clinic shows you the GPU price. Almost none show you the electricity meter running in the background, the Ollama update that breaks your integration every two weeks, or the IT hours that quietly land on your payroll. So here is the position the rest of this piece defends: at low query volume, on-premise is a PDPL compliance decision, not a cost-saving one. This teardown gives you the actual line items behind that call. Hardware, power, software maintenance, and what PDPL compliance really costs to get right.

GPU Options and What Your VRAM Budget Actually Buys

Start with VRAM, because that is what decides which models you can run, not the marketing tier of the card. A single RTX 4090 gives you 24 GB. That comfortably serves an 8B model quantized, with throughput a single clinic will never saturate at triage volumes. Step up to an A6000 and you get 48 GB, which is where you go if you want headroom for a larger model or two models resident at once. An A100 at 80 GB is a 70B-class machine, and for a single clinic it is almost always over-buying.

Now the prices, and these move, so treat them as a snapshot rather than a quote. The RTX 4090 floor in Dubai sits around AED 7,000–11,550 ex-VAT, and you should expect to find it out of stock more often than not. An A6000 runs roughly AED 22,000–23,000 ex-VAT. Those are the gross figures, what you pay at the counter. Hold that thought, because there is a tax section below that pulls real money back off these numbers, and the net cost is the one that belongs in your decision.

Two procurement risks worth naming before you spec an A100. First, it is an export-controlled part. The US AI Diffusion framework was rescinded in May 2025 and a new framework was proposed in March 2026, which means licensing for high-end data-center GPUs is a moving target rather than a settled rule. Build that uncertainty into your timeline. Second, a single GPU is a single point of failure: one card, one machine, one thing to die at the wrong moment. That is a real exposure for a clinic, and I deal with it head-on in the redundancy and TCO sections rather than waving it away here.

DEWA Will Send You a Bill: Running the Power Numbers

People quote Dubai electricity as a flat AED 0.23/kWh and then wonder why the bill is higher. Commercial supply is a four-band slab (AED 0.230, 0.280, 0.320, then 0.380 per kWh as consumption climbs), plus a fuel surcharge of around AED 0.060, a meter charge of roughly AED 20, and 5% VAT on top. Run the arithmetic and the marginal rate lands near AED 0.462/kWh, with a blended rate around AED 0.45/kWh once everything is averaged. That is meaningfully above the 0.40–0.43 the original framing assumed.

So what does the box actually consume? On a realistic single-clinic duty cycle, an inference workstation pulls somewhere in the range of 3,942–6,570 kWh a year. Multiply by the blended rate and you get roughly AED 1,800–3,000 in electricity annually. That is the number people fixate on, and it is the wrong fixation: across a five-year hold, electricity is a low single-digit percentage of total cost of ownership. The hardware and the people maintaining it dwarf it. A full TCO line further down proves that out. For now, just file the power bill under "small, predictable, not the deciding factor."

The Maintenance Calendar Vendors Forget to Show You

The quote ends at the hardware. The work doesn't. An on-premise LLM is a piece of running software, and running software moves under you. Take the inference server: vLLM has shipped a near-continuous release chain that runs through v0.15.0 in January 2026, and it has not stopped since. Ollama is past v0.30.10 across roughly 587 releases, and that pace is confirmed current. Every one of those releases is a change you either adopt deliberately or get dragged into by accident. Pin your versions. Upgrade on your schedule, not on the project's.

Budget the hours honestly. Patching, model version management, monitoring, and the occasional 2 a.m. "why did throughput drop" all land on someone, and that someone needs roughly 3–5 hours a month in steady state. Call it a fraction of an FTE that nonetheless has to exist. A clinic that does not staff this, or contract it, is not running a private LLM. It is running an unpatched server with a model on it, which is a different and worse thing.

The Line Items Vendors Leave Off the BOM: Import Duty, VAT, and What You Actually Get Back

This is where the gross prices above turn into real numbers, and where most quotes go quietly wrong.

Customs first. GPUs and workstations land under HS headings 8471 and 8473.30, and the default is the 5% GCC Common External Tariff. Import into a free zone and duty is suspended while the goods stay inside it. Whether your specific part sits on an exempt list I can't confirm, so plan for 5% and treat a waiver as upside rather than baseline.

VAT is the line that actually moves the decision, and it cuts the other way. The 5% you pay on import is input tax, and for a business making taxable or zero-rated supplies it is recoverable. Qualifying healthcare in the UAE is zero-rated, which is the good case: zero-rated still preserves full input-tax recovery, so the VAT on your GPU comes back. The trap is mixed supply. If part of what the clinic does is VAT-exempt rather than zero-rated, apportionment blocks recovery on the exempt share, and you only claw back the taxable proportion. One more thing people raise and then over-worry: the Capital Assets Scheme only engages at AED 5,000,000 and above, so a clinic LLM build never touches it. Ignore it.

Put it back together on a net basis. A single-4090 build comes out around AED 6,700–11,000 net once recoverable VAT is stripped out. A dual-A6000 build lands near AED 121,000 fully landed, not the AED 140,000 the inflated version of this BOM floats around, with the recoverable VAT already sitting inside that figure. Which brings the article back to its own thesis: once you clear out the tax confusion, the cheap build gets cheaper, and the honest conclusion is that at the 8B tier the decision was never about cost in the first place.

PDPL, NABIDH, and Why On-Premise Is the Compliance Argument, Not the Cost Argument

If cost doesn't decide it, residency does. And the residency rule for a clinic is more specific than "PDPL," which gets cited reflexively and is rarely the operative law.

The operative law is Federal Law No. 2 of 2019, on the use of ICT in health fields. It is the one that actually puts patient health data inside the country, and it has teeth: failing the localisation requirement carries a fine of AED 500,000–700,000. PDPL, which is Federal Decree-Law No. 45 of 2021 administered by the UAE Data Office, sits behind it as reinforcing law for sensitive personal data, but its executive regulations and its penalty schedule have not been issued. So I treat PDPL as the direction of travel and Federal Law No. 2 of 2019 as the enforceable residency rule. You will sometimes see a confident AED 5,000,000 PDPL fine quoted. I have removed it here because I can't verify it against issued regulation, and a number you can't stand behind is worse than no number.

On top of the residency law sits NABIDH, which is mandatory for a DHA licence. NABIDH means your patient databases, your backups, your logs, and your disaster recovery all live on UAE soil. The control detail comes from DHA via HISHD, and the relevant standard is ST-09, issued 02/01/2025 and effective 02/04/2025: encryption, role-based access control, Break-the-Glass emergency access, audit logging, and 25-year retention. Those are not aspirations. They are the spec your deployment has to meet to keep the licence.

Dubai is not the whole country, so state it UAE-wide. Abu Dhabi runs the parallel stack, Malaffi under the DoH, governed by ADHICS v2.0, and at federal level there is Riayati and NUMR under MoHAP. Different names, same shape: patient data stays in-country under a named health authority's controls.

One adjacent regime worth flagging for clinics inside the financial free zone: DIFC Regulation 10 governs autonomous systems, reaches full enforcement in January 2026, and requires certification plus a registered Autonomous Systems Officer. The wording there is settled; if you operate in the DIFC, build for it now.

Now the honest part, because it would be easy to read all of this as "therefore you must buy hardware." It does not say that. Residency says the data stays in the UAE under these controls. It does not say the only compliant place for it is a box in your own server room, which is exactly why the next section exists.

Sovereign Cloud Is the Alternative You Have to Argue Against (Not US-Region Azure)

The strongest case against buying hardware is not "use ChatGPT" and it is not "spin up a US-region Azure tenant." Both of those fail residency on their own terms. The case you actually have to beat is UAE sovereign cloud, and it is a real one.

Two offerings matter. Core42 runs a Sovereign Public Cloud on Azure, mapped to 200+ controls, integrated with NABIDH and Malaffi, and resident in the UAE. Then there is the UAE Sovereign Launchpad, which is e& on the AWS Middle East (UAE) Region with Outposts for on-premises extension; it went live on 3 November 2025 and is explicitly aimed at regulated sectors including healthcare. Both are residency-capable. The honest caveat: formal DHA or DoH processor certification for a given healthcare workload is something I'd confirm in writing before signing, not assume from the marketing.

Now quantify what you give up by not owning the hardware. An 8B model behind a managed API runs somewhere around USD 0.02–0.20 per million tokens. A triage query is roughly 2,000 tokens; at 400 queries a day that is on the order of USD 15–60 a year in inference. Sovereign options carry a premium over commodity API pricing, so treat that range as a lower bound rather than the bill. Even doubled, it is small.

So here is the split, and it preserves the thesis rather than dodging it. Below a few hundred queries a day, sovereign cloud is usually the cheaper compliant answer, full stop. On-premise wins when you have a control or air-gap mandate, when volume is sustained and high, or when the workflow genuinely cannot touch an external link. That is a deliberate data-control decision. It is not a default, and anyone who sells it to a small clinic as a default is selling hardware, not judgment.

Five-Year TCO and the Break-Even Query Volume

A board deck wants one line, five years, all-in. Here it is for both builds.

The 4090 build: net hardware of AED 6,700–11,000, plus AED 9,000–15,000 of electricity over five years, plus AED 15,000–30,000 of maintenance over the same period. Call it roughly AED 30,000–55,000 all-in, which amortises to about AED 6,000–11,000 a year. The dual-A6000 build runs net hardware around AED 115,000, with the same power and maintenance lines on top, and lands well past AED 145,000 over five years.

Set that against the alternative. An in-region 8B API at the volume above costs on the order of USD 15–60 a year at 400 queries a day. State the break-even plainly: at realistic single-clinic volume, the on-prem 4090 never crosses into cheaper-per-query territory against that API. It does not break even on cost. That is not a flaw in the build. It is the answer to the wrong question.

Notice what the TCO line makes obvious. Electricity is the small line; hardware and maintenance dominate. The power-bill fixation that opens most of these conversations is misplaced by an order of magnitude. So size the model first. Prove the volume with live traffic before you buy anything. Then decide on the TCO line and the residency mandate, not on inference cost, which was never going to be the reason.

The Decision Framework: When the Numbers Actually Work

Size the model first. Everything else follows from that one number, and getting it wrong is how clinics end up with an over-bought A100 serving an 8B model at triage volume.

Start there. Validate query volume against real traffic, not a projection, before a single purchase order goes out. Then run it as a three-way decision rather than the false binary of "buy a box or do nothing." Sovereign cloud, meaning Core42 on Azure or e& on AWS, is usually the cheaper compliant answer below a few hundred queries a day, and it doubles as your redundancy and failover story. On-premise earns its place on control, on an air-gap mandate, or on sustained high volume that the per-query economics finally justify. Pick on the consolidated TCO and the residency requirement. Inference cost does not get a vote.

Straight Answers: On-Premise Clinic LLM FAQ

What is the minimum viable spec for an on-premise clinic LLM? A single RTX 4090 with 24 GB of VRAM, running an 8B model quantized. That serves a single clinic's triage volume with throughput to spare, at roughly AED 6,700–11,000 net once recoverable VAT comes off.

Is on-premise cheaper than ChatGPT or sovereign cloud? At single-clinic volume, no. An in-region 8B managed API runs on the order of USD 15–60 a year at 400 queries a day; the on-prem 4090, all-in over five years, sits around AED 30,000–55,000. On-premise is a control decision, not a cost saving.

Where does UAE patient data have to live? In the UAE. The enforceable rule is Federal Law No. 2 of 2019, carrying an AED 500,000–700,000 localisation fine, with NABIDH (DHA/HISHD/ST-09: encryption, RBAC, Break-the-Glass, audit logging, 25-year retention) governing the controls. PDPL reinforces this for sensitive data, though its executive regulations and penalties are not yet issued.

How much power does it draw, and what does DEWA charge? Roughly 3,942–6,570 kWh a year for an inference workstation, at a blended commercial rate near AED 0.45/kWh, so about AED 1,800–3,000 annually. It is a small line in the five-year total.

Can I claim the VAT back? Yes, if the clinic makes zero-rated or taxable supplies. Qualifying healthcare is zero-rated, which preserves full input-tax recovery, so the 5% import VAT comes back. Mixed supply triggers apportionment and you only recover the taxable share. Quote builds net, not gross.

What happens when the GPU fails? You plan for it. Size a UPS at 2,000–3,000 VA (about AED 3,500–6,000), keep a spare GPU on the shelf, and keep a sovereign-cloud endpoint configured for failover. A single card is a single point of failure, and a clinic should treat it as one.

Does any of this apply outside Dubai? Yes. Abu Dhabi runs Malaffi under the DoH with ADHICS v2.0; at federal level there is Riayati and NUMR under MoHAP. The residency principle holds UAE-wide. Only the authority and the framework name change.

Questions about your setup?

We help UAE SMEs build AI systems that are compliant, on-premise, and actually useful. Free initial conversation.