Skip to main content
$ nvidia-smi --query-gpu=name,memory.total --format=csv

We own the silicon your data runs on.

MspCoreX is an AI company that bought its own GPUs. Training and most inference run on dedicated NVIDIA hardware, in our own racks, in the EU — on open-weight models we fine-tune ourselves. The finance copilot and the alert-assist validation pass run on a US provider, and wherever any other feature falls back to an external model, or is served by a specialised one, that provider is named on our sub-processor register.

Four facts, and the machines that back them.

▸ 01 · hardware

The machines are ours.

Not rented capacity. Not a GPU cloud with a friendly region selector. NVIDIA accelerators we bought, racked and operate ourselves — our network, our keys, our maintenance window.

owner: MspCoreX
tenancy: single, ours
region: EU
▸ 02 · models

The weights are open.

We build on open-weight foundation models and fine-tune, distil and evaluate them in-house. No licence that can be revoked. No model that changes underneath your workflows overnight. Where a feature calls a closed API instead, that provider is on our sub-processor register.

base: open weights
tuning: in-house
held by: us
▸ 03 · data

Most models never leave.

Most inference happens on our machines, in the EU. The finance copilot and the alert-assist validation pass run on a US provider, and where any other feature falls back to an external model, or is served by a specialised one, that provider is named on our sub-processor register with its scope and transfer basis — and nothing becomes anyone’s training corpus.

egress: named providers only
ai subprocessors: on the register
your data as training: never
▸ 04 · capacity

Dedicated, not shared.

Four DGX systems carry production, with purpose-built training, inference and AI-technician boxes beside them. No queue behind someone else’s workload, and headroom we add by buying hardware — not by renegotiating a rate limit.

production: 4 × NVIDIA DGX
training ram: 1 TB
rate limits: none

The rack. Seven machines.

Four DGX systems carry production. The training node pairs two RTX PRO 6000 Blackwell cards with a terabyte of system memory — what lets a model hold an entire estate in one context window instead of answering from a keyhole. The inference node carries 512 GB behind its own card for the real-time path. And the AI technician has a node of its own, so a call is answered on the same island the tickets live on. Serial numbers and datacentre attestations available under NDA.

NodeRoleAcceleratorMemoryWhat runs on it
Production clusterDedicated production inference — the fleet that serves the platform4 × NVIDIA DGXGrace-Blackwell, unifiedProduction model serving · redundant across nodes · no shared tenancy
TrainingDistillation, fine-tuning, embeddings, long-context analysis2 × NVIDIA RTX PRO 6000 Blackwell2 × 96 GB GDDR71 TB system RAMTeacher model · embeddings · fine-tuning runs · whole-estate analysis in one context window
InferenceReal-time — triage, summaries, vision/OCRNVIDIA RTX PRO 6000 Blackwell96 GB GDDR7512 GB system RAMOpen-weight multimodal model · sub-second replies on the real-time path
AI technicianAnswers the phone, opens and works tickets — on a node of its ownNVIDIA RTX PRO 6000 BlackwellAMD Threadripper96 GB GDDR7256 GB system RAMPer-tenant numbers · speech-to-text · speech synthesis · ticket creation, updates and — soon — resolution

Where a request actually goes.

   YOUR PEOPLE                      OUR INFRASTRUCTURE · EU · hardware we own
   ───────────                      ─────────────────────────────────────────

  ┌────────────┐   TLS 1.3   ┌─────────────────────────────────────────────┐
  │ ticket     │             │ API                                         │
  │ telemetry  │ ──────────▶ │   ├─▶ PII scrub        email·phone·IBAN·CNP │
  │ session    │             │   ├─▶ injection guard  critical → blocked   │
  └────────────┘             │   ├─▶ tenant scope     bound before read    │
                             │   └─▶ audit event      written, or it 4xx   │
  ┌────────────┐   SIP/RTP   ├─────────────────────────────────────────────┤
  │ phone call │ ──────────▶ │ PBX ──▶ voice orchestrator                  │
  └────────────┘             │        speech-to-text · model · synthesis   │
                             └──────────────────┬──────────────────────────┘
                                                ▼
                              ┌─────────────────────────┐
                              │       AI gateway        │
                              │  routes by feature —    │
                              │  external: on register  │
                              └────────────┬────────────┘
           ┌───────────────────┬───────────┴───────────┬───────────────────┐
           ▼                   ▼                       ▼                   ▼
  ┌──────────────────┐  ┌──────────────────┐  ┌──────────────────┐  ┌──────────────────┐
  │ PRODUCTION       │  │ INFERENCE        │  │ TRAINING         │  │ AI TECHNICIAN    │
  │ 4 × NVIDIA DGX   │  │ RTX PRO 6000     │  │ 2 × RTX PRO 6000 │  │ RTX PRO 6000     │
  │ Grace-Blackwell  │  │ 96 GB GDDR7      │  │ 2 × 96 GB GDDR7  │  │ 96 GB GDDR7      │
  │ unified memory   │  │ 512 GB RAM       │  │ 1 TB RAM         │  │ 256 GB RAM       │
  │                  │  │                  │  │                  │  │ Threadripper     │
  │ serves the fleet │  │ triage · summary │  │ distil·fine-tune │  │ phone · tickets  │
  │ redundant nodes  │  │ voice · OCR      │  │ embeddings·eval  │  │ works the queue  │
  │ no shared tenancy│  │ sub-second       │  │ long contexts    │  │ 24/7, 4 languages│
  └─────────┬────────┘  └────────┬─────────┘  └────────┬─────────┘  └────────┬─────────┘
            └────────────────────┴──────────┬──────────┴─────────────────────┘
                                            ▼
                                   answer + audit row
                                            │
                                            ▼
                                    back to you, or
                                    back down the phone

  ══════════ EGRESS PAST THIS LINE ONLY TO NAMED PROVIDERS ══════════
  primary, fallback and specialised: on the sub-processor register
  no data shared for training · not ours, not anyone's

Most MSP platforms buy AI. We build ours.

A general-purpose model knows what a firewall is. It does not know what your Tuesday looks like. Closing that gap is not prompt engineering — it is a training loop, and ours runs end to end on machines we own. Five stages, each of which can refuse to promote the next.

1 · Capture

Real operational signal from MSP work: tickets that were actually resolved, alerts that turned out to matter, sessions that ended in a fix rather than a shrug. It is consented, anonymised and PII-scrubbed before a single row is written — the scrubber runs in front of the storage layer, not as a cleanup pass behind it. What survives is the shape of the work: symptom, path, resolution. Never the customer behind it.

Scrubbed before storage: email · phone · IBAN · VAT · national ID · card · API key · token

2 · Distil

A large open-weight teacher model on the training node reads the captured work and writes the supervision for the smaller model that serves production — the reasoning, not only the answer. Teaching a fast student what a slow, more capable teacher concluded is how you get production latency without paying for it in quality. The teacher never touches live traffic; its entire job is to write the lesson and hand it over.

Teacher: open-weight, batch only · Student: the model that answers you · Pairs versioned per run

3 · Fine-tune

Adapters trained on our own hardware, on our own corpus, on our own schedule — no queue, no quota, no third party holding the weights hostage. Every run is versioned and every corpus is pinned, so any model that has ever served a request can be rebuilt from the training node months later. When an auditor asks which model answered a specific ticket in March, that is a lookup, not an investigation.

Method: adapter fine-tune · Corpus: pinned per run · Rebuild: reproducible from the pinned set

4 · Evaluate

A candidate does not ship until it beats the model already in production on a frozen evaluation set — frozen meaning it was written before the candidate existed, so nobody can quietly tune toward the exam. The gate is automated and it has teeth: regress on any tracked dimension and the candidate is rejected, with no human in the loop available to argue with it. Most candidates do not pass. That is the gate working, not the gate malfunctioning.

Eval set: frozen before the candidate exists · Regression on any dimension: rejected · Manual override: none

5 · Canary, then promote

New weights serve a slice of real traffic before they serve all of it. Confidence, latency and reversal rate are compared against the outgoing model on the same workload — not against a benchmark, against the actual job. A regression rolls the cohort back on its own, before anyone files a ticket about it. Promotion is simply what happens when a canary stays boring for long enough.

Canary: traffic slice · Watched: confidence · latency · reversal rate · Rollback: automatic, no page

Every stage of that loop happens on hardware in the table above. No stage of it sends your data to anyone, and no stage of it appears on your invoice as a usage line.

Not a voice agent. A technician.

The industry shipped voice bots: a menu with a nicer accent, which takes a message and hands you back the work. We built something else and gave it its own machine. It answers the phone because that is where the work arrives — but answering is the doorway, not the job. It holds a seat in the queue, it has a login, it has permissions, and everything it does lands in the audit trail under its own name.

▸ 01 · answers

Picks up, in their language.

Per-tenant numbers land on its PBX. It greets the caller, recognises them from the number, and speaks the language they opened with — four of them, around the clock.

▸ 02 · understands

Speech, on its own GPU.

Speech-to-text, the language model and the synthesis all run on its RTX PRO 6000. The audio never travels to a speech vendor, and neither does the transcript.

▸ 03 · works the queue

Opens, updates, follows up.

It creates the ticket with recording and transcript attached, reads back the state of an existing one — ETA, last action, who is on it — and schedules the callback it promised.

▸ 04 · knows its limits

Escalates without being asked.

A caller can request a human at any point and gets a warm transfer to the on-call rotation. So does a case the technician judges is not its to take.

It holds an account, not an API key

The distinction is not cosmetic. It authenticates as a service account and goes through the same pipeline a human does: tenant scope, ACL, audit. A ticket it opens at 21:40 carries the same audit chain as one a human opened at noon — same events, same hash chain, same answer when somebody asks six months later who did what and why. You can read its work the way you read a colleague’s, and you can revoke it the way you revoke a colleague’s.

Auth: service account, hashed key · Per-tenant rate limit and IP allowlist · Audit: identical to a human actor

Next: it closes some of them by itself

Answering and logging is the part that works today. The part being built is autonomy — the technician resolving a class of tickets end to end with nobody else touching them: the password reset, the mailbox quota, the printer queue, the licence reassignment. The runbooks already live in the platform and it already holds the permissions to execute them; what it needs is the judgement to know which cases it may take, and the discipline to escalate every case it may not.

We are deliberately unhurried about that line. A technician that closes a ticket it should have escalated costs more trust than the ten it closed correctly earned. So autonomy ships per tenant, per category, opt-in — every autonomous action reversible, every decision in the audit trail, and a human override always one click away.

Status: in development · Rollout: per tenant, per category, opt-in · Every autonomous action reversible

What owning the hardware buys you.

No per-token bill.

AI is in the plan, not on the meter. When a model vendor reprices, your invoice does not move.

No rug-pull.

We hold the weights for the models we host. One serving your workflows today cannot be deprecated out from under them.

Latency you can feel.

Sub-second inference across a LAN hop for everything we host, instead of a transatlantic API call with someone else’s queue in front of it.

An answer for the auditor.

“Which AI subprocessor sees our data?” — for most features, none: they run on GPUs we own, in the EU. The finance copilot and the alert-assist validation pass run on Anthropic; it and the other fallback and specialised-model providers are on our sub-processor register, with scope and transfer basis — one page you can hand over.

Models, data, controls.

Where inference runs
Mostly our own NVIDIA GPUs, in EU datacentres we control. The finance copilot and the alert-assist validation pass run on Anthropic, in the USA.
Model provenance
Open-weight foundation models, fine-tuned in-house. Version pinned per feature.
Third-party model APIs
Anthropic is primary for the finance copilot and the alert-assist validation pass, and the fallback elsewhere. OpenAI and Mistral are tenant-selectable. All are on the sub-processor register.
Your data as training data
Never. Not for us, not for anyone.
Cross-tenant learning
None. One tenant’s data never influences another tenant’s model.
Per-feature opt-out
Available on every AI operation, per tenant.
Audit on every call
Input, output, model, version, latency, confidence, reversed-by — all logged.
Hardware ownership
Purchased, racked and operated by us. Attestations on request under NDA.
// policy  our GPUs · open weights · most models never leave them · every other model provider on the sub-processor register · no cross-tenant learning · full audit on every call · opt-out per operation.