Dev Seth

private models · on-prem · secure by design

Models your clients own, running where their data already lives.

I build specialised models end to end, starting with the data. Structured toward the outcome, fine-tuned and tested against it, then deployed inside the client's own cloud, data centre or air-gapped network.

the process

Seven steps, from data to a model running in your client's environment.

Tools I use across these steps: PyTorch, Unsloth, Modal, MLflow, Docker, Kubernetes, Terraform, AWS SageMaker, AWS Bedrock, Azure AI and LangGraph, with vLLM and SGLang for serving.

01

Define the outcome

Agree what "good" means for the task, build an eval set from real examples, measure the current approach and a frontier API as baselines.

Written success criteria, eval set, baseline scores.

02

Structure the data

Audit sources, clean and dedupe, handle PII, and shape the data into training and eval examples built around the outcome, not just what's available.

Training and eval datasets, data card, PII handling notes.

raw source record

From: sarah.k@example-mail.com
To: support@example-co.com
Subject: Order #48213 arrived damaged

Hi, my order arrived with a cracked screen. Can I get
a replacement? My number is 555-0142 if easier to call.

Thanks,
Sarah Kim
Account: skim_1984

Fictional example, for illustration only.

03

Choose the model specs

Pick the base model family, size, context length and licence against the hardware target, latency budget and data sensitivity; choose the tuning method.

Model spec sheet with trade-offs and hardware estimate.

LoRAQLoRADPOGRPO

Hardware

Latency

Data sensitivity

starting point, not a quote

Medium (7B to 14B params)

QLoRA fine-tune, quantized for serving

04

Fine-tune and iterate

Train, evaluate, read the failures, fix the data, repeat. Automated evals plus LLM-as-judge and human review, every run tracked. Worked example: the Playwright test-generation model.

Model checkpoints, eval reports per iteration, experiment log.

PyTorchUnslothMLflow
05

Compress and harden

Distill and quantize to fit the target hardware, test for regressions, add guardrails. Worked example: the on-device document processing model.

Optimised model, regression suite, guardrail config.

06

Deploy privately

Package and serve inside the client's environment with access control, audit logging, and vLLM/SGLang for high-throughput serving.

Deployment scripts (IaC), serving endpoint, runbook.

DockerKubernetesTerraformvLLMSGLang
07

Monitor and improve

Tracing, drift checks and a retraining cadence tied back to the eval set.

Monitoring dashboard, retraining plan.

why private

Three reasons it's worth doing.

Control

Data, weights and eval sets stay with the client.

Cost

A smaller tuned model can replace frontier API calls on a narrow task, as with the Playwright test-generation model.

Compliance

Runs in the client's region or building, which matters for regulated sectors and data residency requirements.

security by design

Built to hold up under a security review.

◆01

Data stays put

Data never leaves the client environment during training or inference.

◈02

Role-based access

Role-based access to data, weights and endpoints.

◇03

Audit logging

Audit logs on every model call.

◫04

PII handling

PII handling and redaction in the data pipeline.

◆05

Encrypted storage

Encrypted storage for datasets and weights.

◈06

Air-gapped option

Offline model delivery, as built for a national defence programme.

fine-tune, rag, or both?

Not sure which you need? Answer three questions.

question 1 of 3

Is the problem mostly missing knowledge, or wrong behaviour, format or tone?

deployment options

Client VPC, on-prem, on-device, or fully air-gapped.

DataTrainingWeightsInferenceMonitoring

Deployed inside the client's own cloud account and network, isolated from other tenants.

Typical for mid-size enterprise clients on AWS or Azure.

proof

What this looks like, built.

fine-tuned

Playwright test-generation model

A small model fine-tuned to write Playwright browser tests.

More accurate at test generation than frontier models in my benchmarks.

on-device

On-device document model

Processes documents on a phone, with nothing sent to the cloud.

Private by design: the document never leaves the device.

specialised

Decision intelligence engine

A purpose-built model for fast, consistent decisions inside a workflow.

Built for decisions that can't wait on a general-purpose model.

air-gapped

Intelligence correlation model

A specialised model that connects the dots across global information.

Runs entirely inside controlled infrastructure.

See it in production →
forecasting

US home price forecast model

Predicted home price growth across 100 major US markets at 87% accuracy.

Called a 4% rise for 2023 when most expected a fall. Named the most accurate forecast of the year.

See the case studies →

faq

Common questions.

The client. Data, eval sets and weights stay theirs.

The regression suite blocks it before it ships, and the previous version stays live until it's fixed.

Have a use case that can't leave the building?