We own the silicon your data runs on.
MspCoreX is an AI company that bought its own GPUs. Training and most inference run on dedicated NVIDIA hardware, in our own racks, in the EU — on open-weight models we fine-tune ourselves. The finance copilot and the alert-assist validation pass run on a US provider, and wherever any other feature falls back to an external model, or is served by a specialised one, that provider is named on our sub-processor register.
Four facts, and the machines that back them.
The machines are ours.
Not rented capacity. Not a GPU cloud with a friendly region selector. NVIDIA accelerators we bought, racked and operate ourselves — our network, our keys, our maintenance window.
The weights are open.
We build on open-weight foundation models and fine-tune, distil and evaluate them in-house. No licence that can be revoked. No model that changes underneath your workflows overnight. Where a feature calls a closed API instead, that provider is on our sub-processor register.
Most models never leave.
Most inference happens on our machines, in the EU. The finance copilot and the alert-assist validation pass run on a US provider, and where any other feature falls back to an external model, or is served by a specialised one, that provider is named on our sub-processor register with its scope and transfer basis — and nothing becomes anyone’s training corpus.
Dedicated, not shared.
Four DGX systems carry production, with purpose-built training, inference and AI-technician boxes beside them. No queue behind someone else’s workload, and headroom we add by buying hardware — not by renegotiating a rate limit.
The rack. Seven machines.
Four DGX systems carry production. The training node pairs two RTX PRO 6000 Blackwell cards with a terabyte of system memory — what lets a model hold an entire estate in one context window instead of answering from a keyhole. The inference node carries 512 GB behind its own card for the real-time path. And the AI technician has a node of its own, so a call is answered on the same island the tickets live on. Serial numbers and datacentre attestations available under NDA.
| Node | Role | Accelerator | Memory | What runs on it |
|---|---|---|---|---|
| Production cluster | Dedicated production inference — the fleet that serves the platform | 4 × NVIDIA DGX | Grace-Blackwell, unified | Production model serving · redundant across nodes · no shared tenancy |
| Training | Distillation, fine-tuning, embeddings, long-context analysis | 2 × NVIDIA RTX PRO 6000 Blackwell | 2 × 96 GB GDDR71 TB system RAM | Teacher model · embeddings · fine-tuning runs · whole-estate analysis in one context window |
| Inference | Real-time — triage, summaries, vision/OCR | NVIDIA RTX PRO 6000 Blackwell | 96 GB GDDR7512 GB system RAM | Open-weight multimodal model · sub-second replies on the real-time path |
| AI technician | Answers the phone, opens and works tickets — on a node of its own | NVIDIA RTX PRO 6000 BlackwellAMD Threadripper | 96 GB GDDR7256 GB system RAM | Per-tenant numbers · speech-to-text · speech synthesis · ticket creation, updates and — soon — resolution |
Where a request actually goes.
YOUR PEOPLE OUR INFRASTRUCTURE · EU · hardware we own
─────────── ─────────────────────────────────────────
┌────────────┐ TLS 1.3 ┌─────────────────────────────────────────────┐
│ ticket │ │ API │
│ telemetry │ ──────────▶ │ ├─▶ PII scrub email·phone·IBAN·CNP │
│ session │ │ ├─▶ injection guard critical → blocked │
└────────────┘ │ ├─▶ tenant scope bound before read │
│ └─▶ audit event written, or it 4xx │
┌────────────┐ SIP/RTP ├─────────────────────────────────────────────┤
│ phone call │ ──────────▶ │ PBX ──▶ voice orchestrator │
└────────────┘ │ speech-to-text · model · synthesis │
└──────────────────┬──────────────────────────┘
▼
┌─────────────────────────┐
│ AI gateway │
│ routes by feature — │
│ external: on register │
└────────────┬────────────┘
┌───────────────────┬───────────┴───────────┬───────────────────┐
▼ ▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ PRODUCTION │ │ INFERENCE │ │ TRAINING │ │ AI TECHNICIAN │
│ 4 × NVIDIA DGX │ │ RTX PRO 6000 │ │ 2 × RTX PRO 6000 │ │ RTX PRO 6000 │
│ Grace-Blackwell │ │ 96 GB GDDR7 │ │ 2 × 96 GB GDDR7 │ │ 96 GB GDDR7 │
│ unified memory │ │ 512 GB RAM │ │ 1 TB RAM │ │ 256 GB RAM │
│ │ │ │ │ │ │ Threadripper │
│ serves the fleet │ │ triage · summary │ │ distil·fine-tune │ │ phone · tickets │
│ redundant nodes │ │ voice · OCR │ │ embeddings·eval │ │ works the queue │
│ no shared tenancy│ │ sub-second │ │ long contexts │ │ 24/7, 4 languages│
└─────────┬────────┘ └────────┬─────────┘ └────────┬─────────┘ └────────┬─────────┘
└────────────────────┴──────────┬──────────┴─────────────────────┘
▼
answer + audit row
│
▼
back to you, or
back down the phone
══════════ EGRESS PAST THIS LINE ONLY TO NAMED PROVIDERS ══════════
primary, fallback and specialised: on the sub-processor register
no data shared for training · not ours, not anyone'sMost MSP platforms buy AI. We build ours.
A general-purpose model knows what a firewall is. It does not know what your Tuesday looks like. Closing that gap is not prompt engineering — it is a training loop, and ours runs end to end on machines we own. Five stages, each of which can refuse to promote the next.
1 · Capture
Real operational signal from MSP work: tickets that were actually resolved, alerts that turned out to matter, sessions that ended in a fix rather than a shrug. It is consented, anonymised and PII-scrubbed before a single row is written — the scrubber runs in front of the storage layer, not as a cleanup pass behind it. What survives is the shape of the work: symptom, path, resolution. Never the customer behind it.
Scrubbed before storage: email · phone · IBAN · VAT · national ID · card · API key · token
2 · Distil
A large open-weight teacher model on the training node reads the captured work and writes the supervision for the smaller model that serves production — the reasoning, not only the answer. Teaching a fast student what a slow, more capable teacher concluded is how you get production latency without paying for it in quality. The teacher never touches live traffic; its entire job is to write the lesson and hand it over.
Teacher: open-weight, batch only · Student: the model that answers you · Pairs versioned per run
3 · Fine-tune
Adapters trained on our own hardware, on our own corpus, on our own schedule — no queue, no quota, no third party holding the weights hostage. Every run is versioned and every corpus is pinned, so any model that has ever served a request can be rebuilt from the training node months later. When an auditor asks which model answered a specific ticket in March, that is a lookup, not an investigation.
Method: adapter fine-tune · Corpus: pinned per run · Rebuild: reproducible from the pinned set
4 · Evaluate
A candidate does not ship until it beats the model already in production on a frozen evaluation set — frozen meaning it was written before the candidate existed, so nobody can quietly tune toward the exam. The gate is automated and it has teeth: regress on any tracked dimension and the candidate is rejected, with no human in the loop available to argue with it. Most candidates do not pass. That is the gate working, not the gate malfunctioning.
Eval set: frozen before the candidate exists · Regression on any dimension: rejected · Manual override: none
5 · Canary, then promote
New weights serve a slice of real traffic before they serve all of it. Confidence, latency and reversal rate are compared against the outgoing model on the same workload — not against a benchmark, against the actual job. A regression rolls the cohort back on its own, before anyone files a ticket about it. Promotion is simply what happens when a canary stays boring for long enough.
Canary: traffic slice · Watched: confidence · latency · reversal rate · Rollback: automatic, no page
Every stage of that loop happens on hardware in the table above. No stage of it sends your data to anyone, and no stage of it appears on your invoice as a usage line.
Not a voice agent. A technician.
The industry shipped voice bots: a menu with a nicer accent, which takes a message and hands you back the work. We built something else and gave it its own machine. It answers the phone because that is where the work arrives — but answering is the doorway, not the job. It holds a seat in the queue, it has a login, it has permissions, and everything it does lands in the audit trail under its own name.
Picks up, in their language.
Per-tenant numbers land on its PBX. It greets the caller, recognises them from the number, and speaks the language they opened with — four of them, around the clock.
Speech, on its own GPU.
Speech-to-text, the language model and the synthesis all run on its RTX PRO 6000. The audio never travels to a speech vendor, and neither does the transcript.
Opens, updates, follows up.
It creates the ticket with recording and transcript attached, reads back the state of an existing one — ETA, last action, who is on it — and schedules the callback it promised.
Escalates without being asked.
A caller can request a human at any point and gets a warm transfer to the on-call rotation. So does a case the technician judges is not its to take.
It holds an account, not an API key
The distinction is not cosmetic. It authenticates as a service account and goes through the same pipeline a human does: tenant scope, ACL, audit. A ticket it opens at 21:40 carries the same audit chain as one a human opened at noon — same events, same hash chain, same answer when somebody asks six months later who did what and why. You can read its work the way you read a colleague’s, and you can revoke it the way you revoke a colleague’s.
Auth: service account, hashed key · Per-tenant rate limit and IP allowlist · Audit: identical to a human actor
Next: it closes some of them by itself
Answering and logging is the part that works today. The part being built is autonomy — the technician resolving a class of tickets end to end with nobody else touching them: the password reset, the mailbox quota, the printer queue, the licence reassignment. The runbooks already live in the platform and it already holds the permissions to execute them; what it needs is the judgement to know which cases it may take, and the discipline to escalate every case it may not.
We are deliberately unhurried about that line. A technician that closes a ticket it should have escalated costs more trust than the ten it closed correctly earned. So autonomy ships per tenant, per category, opt-in — every autonomous action reversible, every decision in the audit trail, and a human override always one click away.
Status: in development · Rollout: per tenant, per category, opt-in · Every autonomous action reversible
What owning the hardware buys you.
No per-token bill.
AI is in the plan, not on the meter. When a model vendor reprices, your invoice does not move.
No rug-pull.
We hold the weights for the models we host. One serving your workflows today cannot be deprecated out from under them.
Latency you can feel.
Sub-second inference across a LAN hop for everything we host, instead of a transatlantic API call with someone else’s queue in front of it.
An answer for the auditor.
“Which AI subprocessor sees our data?” — for most features, none: they run on GPUs we own, in the EU. The finance copilot and the alert-assist validation pass run on Anthropic; it and the other fallback and specialised-model providers are on our sub-processor register, with scope and transfer basis — one page you can hand over.
Models, data, controls.
- Where inference runs
- Mostly our own NVIDIA GPUs, in EU datacentres we control. The finance copilot and the alert-assist validation pass run on Anthropic, in the USA.
- Model provenance
- Open-weight foundation models, fine-tuned in-house. Version pinned per feature.
- Third-party model APIs
- Anthropic is primary for the finance copilot and the alert-assist validation pass, and the fallback elsewhere. OpenAI and Mistral are tenant-selectable. All are on the sub-processor register.
- Your data as training data
- Never. Not for us, not for anyone.
- Cross-tenant learning
- None. One tenant’s data never influences another tenant’s model.
- Per-feature opt-out
- Available on every AI operation, per tenant.
- Audit on every call
- Input, output, model, version, latency, confidence, reversed-by — all logged.
- Hardware ownership
- Purchased, racked and operated by us. Attestations on request under NDA.