Playwright test-generation model
A small model fine-tuned to write Playwright browser tests.
More accurate at test generation than frontier models in my benchmarks.
private models · on-prem · secure by design
I build specialised models end to end, starting with the data. Structured toward the outcome, fine-tuned and tested against it, then deployed inside the client's own cloud, data centre or air-gapped network.
the process
Tools I use across these steps: PyTorch, Unsloth, Modal, MLflow, Docker, Kubernetes, Terraform, AWS SageMaker, AWS Bedrock, Azure AI and LangGraph, with vLLM and SGLang for serving.
Agree what "good" means for the task, build an eval set from real examples, measure the current approach and a frontier API as baselines.
Written success criteria, eval set, baseline scores.
Audit sources, clean and dedupe, handle PII, and shape the data into training and eval examples built around the outcome, not just what's available.
Training and eval datasets, data card, PII handling notes.
raw source record
From: sarah.k@example-mail.com To: support@example-co.com Subject: Order #48213 arrived damaged Hi, my order arrived with a cracked screen. Can I get a replacement? My number is 555-0142 if easier to call. Thanks, Sarah Kim Account: skim_1984
Fictional example, for illustration only.
Pick the base model family, size, context length and licence against the hardware target, latency budget and data sensitivity; choose the tuning method.
Model spec sheet with trade-offs and hardware estimate.
Hardware
Latency
Data sensitivity
starting point, not a quote
Medium (7B to 14B params)
QLoRA fine-tune, quantized for serving
Train, evaluate, read the failures, fix the data, repeat. Automated evals plus LLM-as-judge and human review, every run tracked. Worked example: the Playwright test-generation model.
Model checkpoints, eval reports per iteration, experiment log.
Distill and quantize to fit the target hardware, test for regressions, add guardrails. Worked example: the on-device document processing model.
Optimised model, regression suite, guardrail config.
Package and serve inside the client's environment with access control, audit logging, and vLLM/SGLang for high-throughput serving.
Deployment scripts (IaC), serving endpoint, runbook.
Tracing, drift checks and a retraining cadence tied back to the eval set.
Monitoring dashboard, retraining plan.
01
what I do
Agree what "good" means for the task, build an eval set from real examples, measure the current approach and a frontier API as baselines.
what you get
Written success criteria, eval set, baseline scores.
02
what I do
Audit sources, clean and dedupe, handle PII, and shape the data into training and eval examples built around the outcome, not just what's available.
what you get
Training and eval datasets, data card, PII handling notes.
raw source record
From: sarah.k@example-mail.com To: support@example-co.com Subject: Order #48213 arrived damaged Hi, my order arrived with a cracked screen. Can I get a replacement? My number is 555-0142 if easier to call. Thanks, Sarah Kim Account: skim_1984
Fictional example, for illustration only.
03
what I do
Pick the base model family, size, context length and licence against the hardware target, latency budget and data sensitivity; choose the tuning method.
what you get
Model spec sheet with trade-offs and hardware estimate.
Hardware
Latency
Data sensitivity
starting point, not a quote
Medium (7B to 14B params)
QLoRA fine-tune, quantized for serving
04
what I do
Train, evaluate, read the failures, fix the data, repeat. Automated evals plus LLM-as-judge and human review, every run tracked. Worked example: the Playwright test-generation model.
what you get
Model checkpoints, eval reports per iteration, experiment log.
05
what I do
Distill and quantize to fit the target hardware, test for regressions, add guardrails. Worked example: the on-device document processing model.
what you get
Optimised model, regression suite, guardrail config.
06
what I do
Package and serve inside the client's environment with access control, audit logging, and vLLM/SGLang for high-throughput serving.
what you get
Deployment scripts (IaC), serving endpoint, runbook.
07
what I do
Tracing, drift checks and a retraining cadence tied back to the eval set.
what you get
Monitoring dashboard, retraining plan.
why private
Data, weights and eval sets stay with the client.
A smaller tuned model can replace frontier API calls on a narrow task, as with the Playwright test-generation model.
Runs in the client's region or building, which matters for regulated sectors and data residency requirements.
security by design
Data never leaves the client environment during training or inference.
Role-based access to data, weights and endpoints.
Audit logs on every model call.
PII handling and redaction in the data pipeline.
Encrypted storage for datasets and weights.
Offline model delivery, as built for a national defence programme.
fine-tune, rag, or both?
question 1 of 3
Is the problem mostly missing knowledge, or wrong behaviour, format or tone?
deployment options
Deployed inside the client's own cloud account and network, isolated from other tenants.
Typical for mid-size enterprise clients on AWS or Azure.
proof
A small model fine-tuned to write Playwright browser tests.
More accurate at test generation than frontier models in my benchmarks.
Processes documents on a phone, with nothing sent to the cloud.
Private by design: the document never leaves the device.
A purpose-built model for fast, consistent decisions inside a workflow.
Built for decisions that can't wait on a general-purpose model.
A specialised model that connects the dots across global information.
Runs entirely inside controlled infrastructure.
See it in production →Predicted home price growth across 100 major US markets at 87% accuracy.
Called a 4% rise for 2023 when most expected a fall. Named the most accurate forecast of the year.
faq
The client. Data, eval sets and weights stay theirs.
The regression suite blocks it before it ships, and the previous version stays live until it's fixed.