Skip to main content
Back to Blog
Small Language ModelsLoRAPEFTOllamaGGUFDockerNext.jsTailwind CSSPostmanMLOps

Fine-Tuning Domain-Specific Small Language Models (SLMs): A Full-Stack LoRA Pipeline with Python, Ollama, and Next.js

A hands-on 2026 guide to fine-tuning domain-specific Small Language Models with LoRA in Python, exporting to GGUF, serving with Ollama in Docker, testing APIs with Postman, and monitoring everything in a Next.js dashboard.

October 6, 202622 min readNiraj Kumar

Not long ago, "adding AI" to an enterprise workflow meant wiring up a big cloud API, crossing your fingers, and then watching the monthly invoice climb. That works fine for prototypes. It gets uncomfortable when a support team is pushing two million tickets a month through the same prompt, your compliance officer is asking where the data goes, and latency spikes every time a provider has a bad afternoon.

In 2026, there's a more grown-up option: train a Small Language Model (SLM) on your own domain, run it on hardware you control, and put a proper API and dashboard around it. The tooling is finally mature enough that one developer can build the whole thing in a week.

This guide walks through a complete, full-stack pipeline. We'll prepare a dataset, fine-tune with LoRA using Python, convert the result to GGUF, serve it with Ollama inside Docker on Linux, validate the REST layer with Postman, and build a real-time observability dashboard with Next.js, React, and Tailwind CSS.

Architecture overview of the full-stack SLM pipeline: dataset, LoRA training, GGUF conversion, Ollama in Docker, Next.js dashboard

Why Small Language Models Make Sense for Enterprise Work

A general-purpose frontier model is like hiring a brilliant polymath to sort your mail. It will do the job, but you're paying polymath rates for a task that needs consistency, not brilliance.

Most enterprise AI workloads are narrow and repetitive:

  • Classifying and routing support tickets
  • Extracting fields from invoices, contracts, or clinical notes
  • Rewriting text into a fixed house style
  • Answering questions over a bounded internal knowledge base
  • Generating structured JSON from messy human input

For these jobs, a fine-tuned SLM has real advantages:

  • Cost: Inference on a 1B to 4B parameter model is cheap, and the cost is mostly fixed (your hardware) rather than variable (per token).
  • Latency: Small models respond fast, and local network hops beat internet round trips.
  • Privacy and sovereignty: Sensitive data never leaves your infrastructure, which makes audits much simpler.
  • Predictability: You pin the model version. Nobody silently swaps it out from under you.
  • Format discipline: Fine-tuning teaches the model your exact output schema, so you rely less on long, fragile prompts.

The trade-off is honest and worth stating: an SLM will not write you a sonnet about Kubernetes and then debug your Rust. It's a specialist. That's the point.

A rough look at the economics

Let's do some back-of-the-envelope math. The numbers below are illustrative assumptions, not quotes. Swap in your own.

Imagine a workload of 2 million requests per month, averaging 750 tokens each (prompt plus completion). That's roughly 1.5 billion tokens a month.

ApproachIllustrative cost modelRough monthly cost
Frontier cloud APIBlended price of about $3 per 1M tokensAbout $4,500
Mid-tier cloud APIBlended price of about $0.60 per 1M tokensAbout $900
Self-hosted SLMOne GPU server, amortized hardware plus powerA few hundred dollars, mostly fixed

The self-hosted number barely moves as traffic grows, until you saturate the GPU. The API numbers scale linearly. That's the whole story: the more steady, high-volume, and narrow your workload, the better the case for owning the model. If your traffic is tiny or unpredictable, stay on the API and be happy about it.

The Pipeline at a Glance

Before diving into code, here's the flow we'll build:

  1. Prepare data: Clean, deduplicate, and format examples as chat-style JSONL.
  2. Fine-tune: Train LoRA adapters on a base SLM using PEFT and TRL.
  3. Evaluate: Measure format validity and accuracy on a held-out set.
  4. Merge and convert: Fold the adapter into the base model, export to GGUF, and quantize.
  5. Serve: Package it with an Ollama Modelfile and run it in Docker.
  6. Validate: Test the REST contract with Postman.
  7. Observe: Build a Next.js dashboard showing latency, throughput, and error rates.

Our running example is an IT support ticket triage assistant. It reads a raw ticket and returns strict JSON: a category, a priority, and a one-line summary. It's small, realistic, and easy to evaluate.

Step 1: Dataset Preparation (Where Most Projects Are Won or Lost)

If you remember one thing from this article, let it be this: your model can only be as good as your data. LoRA is forgiving about compute. It is not forgiving about messy labels.

Pick a format that matches inference

Use the same chat structure at training time that you'll use at inference time. A JSONL file where each line contains a messages array works with most modern trainers:

{"messages": [
  {"role": "system", "content": "You are a ticket triage assistant. Reply with JSON only."},
  {"role": "user", "content": "VPN keeps disconnecting every 10 minutes since the update. Whole sales team affected."},
  {"role": "assistant", "content": "{\"category\": \"network\", \"priority\": \"high\", \"summary\": \"Recurring VPN drops affecting the sales team after an update.\"}"}
]}

Clean, dedupe, split

Here's a practical preparation script. It validates the assistant output, removes near-duplicates, and creates train and validation splits.

# prepare_dataset.py
import json
import random
import hashlib
from pathlib import Path

ALLOWED_CATEGORIES = {"network", "hardware", "software", "access", "other"}
ALLOWED_PRIORITIES = {"low", "medium", "high", "critical"}

def is_valid(example: dict) -> bool:
    try:
        assistant = example["messages"][-1]["content"]
        payload = json.loads(assistant)
        return (
            payload["category"] in ALLOWED_CATEGORIES
            and payload["priority"] in ALLOWED_PRIORITIES
            and isinstance(payload["summary"], str)
            and 5 < len(payload["summary"]) < 200
        )
    except (KeyError, IndexError, json.JSONDecodeError, TypeError):
        return False

def fingerprint(example: dict) -> str:
    user_text = example["messages"][1]["content"].lower().strip()
    return hashlib.sha1(user_text.encode("utf-8")).hexdigest()

def main(src="raw_tickets.jsonl", out_dir="data", val_ratio=0.1, seed=42):
    random.seed(seed)
    seen, kept = set(), []

    for line in Path(src).read_text(encoding="utf-8").splitlines():
        ex = json.loads(line)
        fp = fingerprint(ex)
        if fp in seen or not is_valid(ex):
            continue
        seen.add(fp)
        kept.append(ex)

    random.shuffle(kept)
    split = int(len(kept) * (1 - val_ratio))
    train, val = kept[:split], kept[split:]

    Path(out_dir).mkdir(exist_ok=True)
    for name, rows in (("train", train), ("val", val)):
        with open(f"{out_dir}/{name}.jsonl", "w", encoding="utf-8") as f:
            for row in rows:
                f.write(json.dumps(row, ensure_ascii=False) + "\n")

    print(f"kept={len(kept)} train={len(train)} val={len(val)}")

if __name__ == "__main__":
    main()

A few hard-won lessons about data:

  • Strip personal data first. Names, emails, and account numbers have a way of being memorized.
  • Balance your classes. If 80% of tickets are "software," the model will happily guess "software" forever.
  • Include hard and ugly examples. Typos, mixed languages, and vague tickets are what production actually looks like.
  • Keep a golden set. Hold out a few hundred reviewed examples that never touch training. This is your truth source.

Step 2: Fine-Tuning with PEFT and LoRA

What LoRA actually does

Full fine-tuning updates every weight in the model, which is slow and memory-hungry. LoRA (Low-Rank Adaptation) freezes the original weights and injects small trainable matrices into selected layers. Instead of learning a giant update matrix, it learns two skinny ones whose product approximates the change. You end up training often well under 1% of the parameters.

QLoRA goes one step further by loading the frozen base model in 4-bit precision, so the memory footprint shrinks dramatically while the adapters still train in higher precision.

The knobs you'll tune most often:

  • r (rank): The capacity of the adapter. 8 to 32 covers most cases.
  • lora_alpha: A scaling factor, commonly set to 1x to 2x the rank.
  • target_modules: Which layers get adapters. Targeting all linear projection layers usually performs best.
  • lora_dropout: Light regularization, typically 0.05.

The training script

We'll use Qwen/Qwen2.5-1.5B-Instruct as the base because it's small, permissively licensed, and punches above its weight. You can swap in another instruct-tuned SLM, such as a Llama 3.2 or Phi variant, as long as you check its license.

Install dependencies and pin the versions you tested, since the Hugging Face ecosystem moves quickly:

python -m venv .venv && source .venv/bin/activate
pip install torch transformers peft trl datasets accelerate bitsandbytes
pip freeze > requirements.lock.txt
# train_lora.py
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

BASE_MODEL = "Qwen/Qwen2.5-1.5B-Instruct"
OUTPUT_DIR = "outputs/ticket-triage-lora"

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL)
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    quantization_config=bnb_config,
    device_map="auto",
)

lora_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
    target_modules=[
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj",
    ],
)

dataset = load_dataset(
    "json",
    data_files={"train": "data/train.jsonl", "validation": "data/val.jsonl"},
)

training_args = SFTConfig(
    output_dir=OUTPUT_DIR,
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    logging_steps=10,
    eval_strategy="steps",
    eval_steps=100,
    save_strategy="steps",
    save_steps=100,
    save_total_limit=2,
    bf16=True,
    gradient_checkpointing=True,
    max_length=1024,
    report_to="none",
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset["train"],
    eval_dataset=dataset["validation"],
    peft_config=lora_config,
    processing_class=tokenizer,
)

trainer.train()
trainer.save_model(OUTPUT_DIR)
tokenizer.save_pretrained(OUTPUT_DIR)

Run it with python train_lora.py and watch the validation loss. If training loss keeps dropping while validation loss climbs, you're overfitting. Reduce epochs, lower the rank, or add more varied data.

Merge the adapter into the base model

Ollama needs a standalone model, so fold the LoRA weights back into the base. Load the base in 16-bit (not 4-bit) for a clean merge:

# merge_adapter.py
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE_MODEL = "Qwen/Qwen2.5-1.5B-Instruct"
ADAPTER_DIR = "outputs/ticket-triage-lora"
MERGED_DIR = "outputs/ticket-triage-merged"

base = AutoModelForCausalLM.from_pretrained(BASE_MODEL, torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, ADAPTER_DIR)
merged = model.merge_and_unload()

merged.save_pretrained(MERGED_DIR, safe_serialization=True)
AutoTokenizer.from_pretrained(ADAPTER_DIR).save_pretrained(MERGED_DIR)
print("Merged model saved to", MERGED_DIR)

Diagram showing LoRA low-rank adapter matrices A and B injected beside frozen base weights, then merged for export

Step 3: Evaluate Before You Ship

Never skip this. A model that "feels good" in a few manual tries can still fail 15% of the time in production. Build a small evaluation script that checks two things on your golden set: is the output valid JSON that matches the schema, and is the label correct.

# evaluate.py
import json
import requests

OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "ticket-triage"

def predict(ticket: str) -> dict | None:
    r = requests.post(OLLAMA_URL, json={
        "model": MODEL,
        "stream": False,
        "format": "json",
        "options": {"temperature": 0.1},
        "messages": [
            {"role": "system", "content": "You are a ticket triage assistant. Reply with JSON only."},
            {"role": "user", "content": ticket},
        ],
    }, timeout=60)
    r.raise_for_status()
    try:
        return json.loads(r.json()["message"]["content"])
    except json.JSONDecodeError:
        return None

def main(path="data/val.jsonl"):
    total = valid = correct = 0
    for line in open(path, encoding="utf-8"):
        ex = json.loads(line)
        gold = json.loads(ex["messages"][-1]["content"])
        pred = predict(ex["messages"][1]["content"])
        total += 1
        if pred and {"category", "priority", "summary"} <= pred.keys():
            valid += 1
            if pred["category"] == gold["category"]:
                correct += 1
    print(f"format_valid={valid/total:.1%}  category_accuracy={correct/total:.1%}  n={total}")

if __name__ == "__main__":
    main()

Track these numbers for the base model (with a good prompt) and for your fine-tuned model. The gap between them is the business case for the whole project.

Step 4: Convert to GGUF and Quantize

Ollama runs models through llama.cpp, which uses the GGUF format. Conversion is a two-step process: export to a full-precision GGUF, then quantize.

# One-time setup
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
pip install -r requirements.txt
cmake -B build && cmake --build build --config Release -j

# 1) Convert the merged Hugging Face model to an f16 GGUF
python convert_hf_to_gguf.py ../outputs/ticket-triage-merged \
  --outfile ../outputs/ticket-triage-f16.gguf \
  --outtype f16

# 2) Quantize to a compact 4-bit variant
./build/bin/llama-quantize \
  ../outputs/ticket-triage-f16.gguf \
  ../outputs/ticket-triage-q4_k_m.gguf \
  Q4_K_M

How do you pick a quantization level? Here's a simple mental model:

  • Q4_K_M: The popular default. Small, fast, and quality is usually close to the original for focused tasks.
  • Q5_K_M: A bit larger, a bit more faithful. Good when your task is sensitive to subtle wording.
  • Q8_0: Near-lossless, but larger and slower. Useful as a quality reference.

Because your use case is narrow, run your evaluation script against more than one quantization level. You may find Q4_K_M is more than good enough, and that saves real memory and money at scale.

Step 5: Package with Ollama and Run in Docker on Linux

An Ollama Modelfile describes how to run your GGUF: the template, default parameters, and system prompt. Qwen-family models use a ChatML-style template:

# Modelfile
FROM ./ticket-triage-q4_k_m.gguf

TEMPLATE """{{- if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""

SYSTEM """You are a ticket triage assistant. Reply with JSON only."""

PARAMETER temperature 0.1
PARAMETER num_ctx 2048
PARAMETER stop "<|im_end|>"

Always double-check that the template matches the one the base model was trained with. A mismatch here is the most common reason a great fine-tune behaves strangely after deployment.

Now containerize the runtime. On a Linux host with an NVIDIA GPU and the NVIDIA Container Toolkit installed, a docker-compose.yml like this keeps everything reproducible:

# docker-compose.yml
services:
  ollama:
    image: ollama/ollama:latest
    container_name: slm-ollama
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
      - ./models:/models:ro
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "ollama", "list"]
      interval: 15s
      timeout: 5s
      retries: 5

  dashboard:
    build: ./dashboard
    container_name: slm-dashboard
    environment:
      - OLLAMA_URL=http://ollama:11434
      - OLLAMA_MODEL=ticket-triage
    ports:
      - "3000:3000"
    depends_on:
      ollama:
        condition: service_healthy
    restart: unless-stopped

volumes:
  ollama_data:

Place your Modelfile and the .gguf file in ./models, bring the stack up, and register the model once:

docker compose up -d ollama
docker exec -it slm-ollama sh -c "cd /models && ollama create ticket-triage -f Modelfile"
docker exec -it slm-ollama ollama run ticket-triage "Printer on floor 3 won't connect to Wi-Fi."

No GPU? Remove the deploy block. Quantized SLMs run acceptably on modern CPUs for lower-volume workloads, just with fewer tokens per second.

Step 6: Validate the REST API with Postman

Ollama exposes a REST API, and our Next.js app will wrap it with a cleaner, business-specific endpoint. Both deserve contract tests.

Create a Postman collection with a POST {{baseUrl}}/api/triage request:

{
  "ticket": "Cannot log in to the HR portal after password reset. Payroll closes today."
}

Then add this script on the Tests tab:

pm.test("status is 200", () => pm.response.to.have.status(200));

pm.test("responds within latency budget", () => {
  pm.expect(pm.response.responseTime).to.be.below(3000);
});

const body = pm.response.json();

pm.test("has a valid triage object", () => {
  pm.expect(body).to.have.property("triage");
  pm.expect(["network", "hardware", "software", "access", "other"])
    .to.include(body.triage.category);
  pm.expect(["low", "medium", "high", "critical"])
    .to.include(body.triage.priority);
  pm.expect(body.triage.summary).to.be.a("string").and.not.empty;
});

pm.test("exposes performance metadata", () => {
  pm.expect(body.metrics).to.have.property("tokensPerSecond");
  pm.expect(body.metrics.latencyMs).to.be.a("number");
});

Add a second request that sends an empty or oversized ticket and asserts a 400 response. Negative tests catch the ugly bugs that happy paths never will. Once the collection feels solid, run it in CI with Newman:

npx newman run slm-triage.postman_collection.json \
  --env-var baseUrl=http://localhost:3000

Step 7: Build the Observability Dashboard with Next.js, React, and Tailwind

A model you can't observe is a model you can't trust. We'll build a small dashboard that answers the questions on-call engineers actually ask: Is it up? Is it fast? Is it returning valid output?

Mockup of the Next.js SLM observability dashboard showing latency percentiles, tokens per second, and success rate cards

A lightweight metrics store

Ollama's non-streaming responses include helpful timing fields, such as eval_count (tokens generated) and eval_duration (in nanoseconds). We can derive tokens per second from them.

// lib/metrics.ts
export type Sample = {
  at: number;
  latencyMs: number;
  tokensPerSecond: number;
  ok: boolean;
};

const MAX_SAMPLES = 500;
const samples: Sample[] = [];

export function record(sample: Sample) {
  samples.push(sample);
  if (samples.length > MAX_SAMPLES) samples.shift();
}

function percentile(sorted: number[], p: number) {
  if (sorted.length === 0) return 0;
  const idx = Math.min(sorted.length - 1, Math.floor((p / 100) * sorted.length));
  return sorted[idx];
}

export function summarize() {
  const latencies = samples.map((s) => s.latencyMs).sort((a, b) => a - b);
  const okCount = samples.filter((s) => s.ok).length;
  const avgTps =
    samples.length === 0
      ? 0
      : samples.reduce((sum, s) => sum + s.tokensPerSecond, 0) / samples.length;

  return {
    total: samples.length,
    successRate: samples.length ? okCount / samples.length : 1,
    p50: percentile(latencies, 50),
    p95: percentile(latencies, 95),
    avgTokensPerSecond: Number(avgTps.toFixed(1)),
    recent: samples.slice(-30),
  };
}

This in-memory store is perfect for a single-instance demo. For multi-instance production setups, push the same data to Prometheus, OpenTelemetry, or a time-series database instead.

The triage API route

// app/api/triage/route.ts
import { NextResponse } from "next/server";
import { z } from "zod";
import { record } from "@/lib/metrics";

const OLLAMA_URL = process.env.OLLAMA_URL ?? "http://localhost:11434";
const MODEL = process.env.OLLAMA_MODEL ?? "ticket-triage";

const BodySchema = z.object({ ticket: z.string().min(5).max(4000) });

const TriageSchema = z.object({
  category: z.enum(["network", "hardware", "software", "access", "other"]),
  priority: z.enum(["low", "medium", "high", "critical"]),
  summary: z.string().min(1),
});

export async function POST(req: Request) {
  const parsed = BodySchema.safeParse(await req.json().catch(() => null));
  if (!parsed.success) {
    return NextResponse.json({ error: "Invalid request body" }, { status: 400 });
  }

  const started = performance.now();

  try {
    const res = await fetch(`${OLLAMA_URL}/api/chat`, {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({
        model: MODEL,
        stream: false,
        format: "json",
        options: { temperature: 0.1 },
        messages: [
          { role: "system", content: "You are a ticket triage assistant. Reply with JSON only." },
          { role: "user", content: parsed.data.ticket },
        ],
      }),
      signal: AbortSignal.timeout(30_000),
    });

    if (!res.ok) throw new Error(`Ollama responded with ${res.status}`);

    const data = await res.json();
    const latencyMs = Math.round(performance.now() - started);
    const tokensPerSecond =
      data.eval_duration > 0 ? data.eval_count / (data.eval_duration / 1e9) : 0;

    const triage = TriageSchema.safeParse(JSON.parse(data.message.content));
    record({ at: Date.now(), latencyMs, tokensPerSecond, ok: triage.success });

    if (!triage.success) {
      return NextResponse.json(
        { error: "Model returned an invalid structure" },
        { status: 502 },
      );
    }

    return NextResponse.json({
      triage: triage.data,
      metrics: { latencyMs, tokensPerSecond: Number(tokensPerSecond.toFixed(1)) },
    });
  } catch (err) {
    record({
      at: Date.now(),
      latencyMs: Math.round(performance.now() - started),
      tokensPerSecond: 0,
      ok: false,
    });
    return NextResponse.json({ error: "Inference failed" }, { status: 503 });
  }
}

Notice that we validate the model's output with Zod before trusting it. Even a well-tuned SLM can occasionally produce something odd, and your API should never pass that along blindly.

// app/api/metrics/route.ts
import { NextResponse } from "next/server";
import { summarize } from "@/lib/metrics";

export const dynamic = "force-dynamic";

export function GET() {
  return NextResponse.json(summarize());
}

The real-time dashboard component

// app/dashboard/MetricsPanel.tsx
"use client";

import { useEffect, useState } from "react";

type Summary = {
  total: number;
  successRate: number;
  p50: number;
  p95: number;
  avgTokensPerSecond: number;
  recent: { at: number; latencyMs: number; ok: boolean }[];
};

function StatCard({ label, value, hint }: { label: string; value: string; hint?: string }) {
  return (
    <div className="rounded-2xl border border-slate-200 bg-white p-5 shadow-sm dark:border-slate-800 dark:bg-slate-900">
      <p className="text-sm font-medium text-slate-500">{label}</p>
      <p className="mt-2 text-3xl font-semibold tracking-tight text-slate-900 dark:text-slate-50">
        {value}
      </p>
      {hint && <p className="mt-1 text-xs text-slate-400">{hint}</p>}
    </div>
  );
}

export default function MetricsPanel() {
  const [data, setData] = useState<Summary | null>(null);
  const [error, setError] = useState(false);

  useEffect(() => {
    let active = true;

    async function load() {
      try {
        const res = await fetch("/api/metrics", { cache: "no-store" });
        if (!res.ok) throw new Error("bad status");
        const json = (await res.json()) as Summary;
        if (active) {
          setData(json);
          setError(false);
        }
      } catch {
        if (active) setError(true);
      }
    }

    load();
    const id = setInterval(load, 2000);
    return () => {
      active = false;
      clearInterval(id);
    };
  }, []);

  if (!data) {
    return <p className="p-6 text-slate-500">Loading live metrics…</p>;
  }

  const maxLatency = Math.max(1, ...data.recent.map((s) => s.latencyMs));

  return (
    <section className="space-y-6 p-6">
      {error && (
        <p className="rounded-lg bg-amber-50 px-4 py-2 text-sm text-amber-800">
          Connection hiccup. Showing the last known values.
        </p>
      )}

      <div className="grid gap-4 sm:grid-cols-2 lg:grid-cols-4">
        <StatCard label="Requests tracked" value={String(data.total)} />
        <StatCard label="Success rate" value={`${(data.successRate * 100).toFixed(1)}%`} />
        <StatCard label="Latency p50 / p95" value={`${data.p50} / ${data.p95} ms`} />
        <StatCard label="Tokens per second" value={String(data.avgTokensPerSecond)} hint="Average across recent requests" />
      </div>

      <div className="rounded-2xl border border-slate-200 bg-white p-5 dark:border-slate-800 dark:bg-slate-900">
        <h2 className="mb-4 text-sm font-semibold text-slate-700 dark:text-slate-200">
          Recent request latency
        </h2>
        <div className="flex h-32 items-end gap-1">
          {data.recent.map((s) => (
            <div
              key={s.at}
              title={`${s.latencyMs} ms`}
              style={{ height: `${(s.latencyMs / maxLatency) * 100}%` }}
              className={`w-full rounded-t ${s.ok ? "bg-emerald-500" : "bg-rose-500"}`}
            />
          ))}
        </div>
      </div>
    </section>
  );
}

Polling every two seconds is simple, reliable, and perfectly adequate here. If you later need sub-second updates across many viewers, switch the metrics route to Server-Sent Events.

Finally, a small multi-stage Dockerfile for the dashboard (enable output: "standalone" in next.config.ts first):

# dashboard/Dockerfile
FROM node:22-alpine AS deps
WORKDIR /app
COPY package*.json ./
RUN npm ci

FROM node:22-alpine AS builder
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY . .
RUN npm run build

FROM node:22-alpine AS runner
WORKDIR /app
ENV NODE_ENV=production
COPY --from=builder /app/.next/standalone ./
COPY --from=builder /app/.next/static ./.next/static
COPY --from=builder /app/public ./public
EXPOSE 3000
CMD ["node", "server.js"]

Real-World Example: Putting It All Together

Picture a mid-sized logistics company with an internal IT desk handling around 60,000 tickets a month. Their first version used a hosted frontier model with a long prompt full of examples. It worked, but the prompt cost added up, latency varied, and security wasn't thrilled about ticket text containing employee details leaving the network.

They collected about 4,000 historical tickets, had two senior agents relabel a few hundred for quality, and followed the pipeline above. After fine-tuning a 1.5B model, the typical outcomes teams report in this kind of setup look like this:

  • Output format validity jumps close to 100% because the schema is baked into the weights, not begged for in the prompt.
  • Prompt length shrinks dramatically since the long instructions and few-shot examples are no longer needed.
  • Median latency drops because the model is small and sits one network hop away.
  • The monthly bill turns into a mostly fixed infrastructure cost.

Treat those as directional, not guaranteed. Your own evaluation set is the only number that matters.

Best Practices

  • Start with the smallest model that could work. Move up in size only when your evaluation proves you need to.
  • Version everything. Datasets, adapters, GGUF files, Modelfiles, and evaluation results should all be traceable to a Git commit or a registry tag.
  • Prefer instruct-tuned bases for chat-style tasks. They adapt faster and need less data.
  • Keep the prompt at inference identical to the format you trained on, including the system message.
  • Use low temperature (0.0 to 0.2) for extraction and classification tasks.
  • Constrain output. Ollama's format: "json" option and schema validation on your side give you two layers of protection.
  • Run evaluations in CI. Every new adapter should beat the previous one on your golden set before it ships.
  • Plan for rollback. Keep the previous model tag registered in Ollama so you can switch back with one config change.
  • Secure the runtime. Don't expose port 11434 to the internet. Put it on a private Docker network and let only your API layer talk to it.

Common Mistakes to Avoid

  1. Training on unreviewed, auto-generated labels. If a larger model labeled your data, spot-check it. Errors in labels become errors in your model.
  2. Leaking validation data into training. Duplicates across splits make your metrics look fantastic and your production results look sad.
  3. Mismatched chat templates. Training with one format and serving with another silently degrades quality.
  4. Over-training. More epochs aren't always better. Watch validation loss, not just training loss.
  5. Judging by vibes. Five impressive demo prompts are not an evaluation.
  6. Quantizing blindly. Always re-run your evaluation after quantization. Occasionally a particular task is more sensitive than expected.
  7. Ignoring context length. Setting num_ctx too high wastes memory, and too low truncates inputs. Match it to your real ticket sizes.
  8. No observability. Without latency, error, and throughput metrics, you'll find out about problems from angry users.
  9. Forgetting the license. Check that your base model's license permits your commercial use case.

🚀 Pro Tips

  • Mix in a little general data (a small percentage of generic instruction examples) to reduce "catastrophic forgetting" if you need the model to stay conversational.
  • Use rejection sampling for data growth. Generate extra training examples with a stronger model, then keep only the ones that pass your validators and a human spot check.
  • Keep adapters separate during experimentation. Train several LoRA adapters for different tasks on the same base and compare them before you merge anything.
  • Warm the model on startup. Send a dummy request when the container boots so the first real user doesn't pay the model-load penalty. Ollama's keep_alive setting also helps keep the model resident in memory.
  • Log prompts and outputs carefully. Captured production failures are the best raw material for your next training round, as long as you scrub sensitive fields first.
  • Benchmark on your target hardware. Tokens per second on a developer laptop says very little about your production server.
  • Add a confidence fallback. If the model returns an invalid structure or an "other" category too often, route the ticket to a human or a larger model rather than guessing.

📌 Key Takeaways

  • Small Language Models are the pragmatic choice for narrow, high-volume enterprise tasks, offering lower cost, lower latency, and stronger data control.
  • Dataset quality is the biggest lever. Clean, deduplicated, well-labeled examples beat sheer volume every time.
  • LoRA and QLoRA make fine-tuning accessible on a single GPU by training small adapters instead of the full model.
  • Merge your adapter, convert to GGUF with llama.cpp, and quantize to deploy through Ollama on modest hardware.
  • Docker gives you a reproducible, portable runtime, and Postman contract tests keep your REST layer honest.
  • A Next.js and Tailwind observability dashboard turns your model from a black box into something your team can operate confidently.

Conclusion

The shift from rented intelligence to owned intelligence isn't about ideology. It's about fit. When the task is narrow and the volume is steady, a small, well-trained model running on your own hardware is often faster, cheaper, and easier to govern than a giant general-purpose API.

The good news is that the path is now well-paved. A clean dataset, a LoRA training run, a GGUF export, an Ollama container, a handful of Postman tests, and a simple dashboard together form a pipeline that one engineer can stand up in days, not quarters. Start small. Pick one painful workflow, measure the baseline, fine-tune, and let the numbers speak for you.

If you build this yourself, I'd love to hear what you learn, especially about the data. It's always the data.

References

Frequently asked questions

What is a Small Language Model (SLM) and how is it different from an LLM?

An SLM is a language model that typically has somewhere between a few hundred million and roughly 10 billion parameters. It is small enough to run on a single GPU or even a CPU. Unlike giant general-purpose LLMs, SLMs shine when you fine-tune them for one well-defined domain, such as ticket triage, contract clause extraction, or internal knowledge Q&A.

How much data do I need to fine-tune an SLM with LoRA?

For a narrow task with a fixed output format, a clean set of 1,000 to 5,000 high-quality examples is often enough to see a big jump over the base model. Quality matters far more than quantity. A few hundred carefully reviewed examples can beat tens of thousands of noisy ones.

Can I fine-tune with LoRA on a consumer GPU?

Yes. With QLoRA (4-bit quantized base weights plus LoRA adapters), a 1.5B to 3B parameter model can usually be trained on a GPU with 12 to 16 GB of VRAM. Larger models need more memory, but the approach scales down well for experimentation.

Why convert the model to GGUF for Ollama?

Ollama runs models through llama.cpp, which uses the GGUF file format. GGUF packs the weights, tokenizer, and metadata into one file and supports quantization, so your model loads fast and fits in far less memory than the original Hugging Face checkpoint.

Is a self-hosted fine-tuned SLM actually cheaper than a cloud API?

It depends on volume. At low or spiky traffic, pay-per-token APIs are hard to beat. At steady, high-volume workloads with a narrow task, the fixed cost of a GPU server is usually much lower than per-token pricing. Calculate your own break-even point using your real token counts and hardware costs.

Do I still need Postman if I have unit tests?

Unit tests verify your code. Postman collections verify the contract your API presents to real clients: status codes, JSON shapes, latency budgets, and error handling. They are also easy to share with frontend developers and to run in CI with Newman.

Discussion

All Articles
Small Language ModelsLoRAPEFTOllamaGGUFDockerNext.jsTailwind CSSPostmanMLOps

Written by

Niraj Kumar

Software Developer — building scalable systems for businesses.

Building this for real? LangChain Developer Services — Production LangChain & LangGraph systems for RAG, agents, and AI workflows.