MoE from the brick up

A hands-on course on mixture of experts models: how tokens, experts and routers fit together, what quantization does to memory, how to host a 300B+ model on Mac Studios or H100s, how training and LoRA fine-tuning work, and how to build an RCM agent team on a realistic budget, plus agent memory, reinforcement learning and recursive self-improvement. Finish with a 35-question quiz.

Idle expert Routed (10) Shared (1)

These are the 512 experts in one MoE layer of Qwen3.5-397B. Each token wakes up only 11 of them.

1

A token is a row of 4,096 numbers

Text is chopped into tokens: whole words or pieces of words. "Credentialing" might become three tokens. Each token gets an ID from the vocabulary, and that ID is looked up in an embedding table, which returns a list of decimal numbers. From then on, that list is the token.

The word credentialing split into three tokens, each looked up in the embedding table and turned into a vector "Credentialing" Credentialing ID 3,421ID 18,207ID 287 Embedding table lookup Each token becomes 4,096 numbers (16 shown each)
Token IDs shown are illustrative. Real vectors are 4,096 numbers wide; 16 are drawn per token.

The length of that list is the model's hidden dimension, written d_model. Every token has the same width, at every layer, whether it is "the", a comma or "cardiology". A long word does not get a wider vector. It gets split into more tokens.

ModelHidden width (d_model)
Mixtral 8x7B, Mistral 7B4,096
Qwen3.5-397B-A17B (our main example)4,096
Mixtral 8x22B6,144
DeepSeek V3, Kimi K27,168

The 4,096 values are numbers, not bits. Each number is stored using some number of bits: 16 in BF16, 8 in FP8, about 4 when quantized. One token at BF16 is 4,096 × 16 bits = 8 KB. Module 4 builds on this.

2

An expert is a brick with a fixed stud pattern

Every expert is a small feed-forward network. Its input and output must be exactly d_model wide so it snaps onto the rest of the model. Inside, it expands the token to a wider working space, filters it, and compresses it back. That inner width is d_ff, the intermediate size.

token (4,096) ─▶ gate proj  ─┐
token (4,096) ─▶ up proj    ─┴▶ SwiGLU filter (d_ff) ─▶ down proj ─▶ out (4,096)

Three matrices, so parameters per expert per layer = 3 × d_model × d_ff.

Coarse experts versus fine-grained experts

Mixtral uses a few big bricks. Newer models use many small ones, which lets the router mix specialists more precisely.

Mixtral expert widens from 4,096 to 14,336 and back; Qwen3.5 expert narrows from 4,096 to 1,024 and back Mixtral expertQwen3.5 expert 4,09614,3364,096 4,0961,0244,096 gate + updown 8 per layer, 2 active~176M params each 512 per layer, 11 active~12.6M params each
Same 4,096 studs top and bottom. The inner width is a design choice.
Mixtral 8x7BQwen3.5-397B
Experts per layer8512
Active per token210 routed + 1 shared
Expert inner width (d_ff)14,3361,024
Params per expert per layer~176M~12.6M
Layers3260

Check the math yourself

Qwen3.5 expert: 3 × 4,096 × 1,024 = 12.58M params
512 experts × 12.58M = 6.44B per layer
6.44B × 60 layers ≈ 386B in experts
+ attention, embeddings, vision encoder ≈ 397B total

Active per token: 11 experts × 12.58M × 60 ≈ 8.3B
+ attention and embeddings ≈ 17B active

That is the whole trick of MoE. You store 397B parameters, but each token only touches about 17B of them. Memory scales with the total. Compute per token scales with the active count.

A shared expert is one brick every token always passes through. It holds common knowledge (grammar, general facts) so the routed experts can specialize.

3

The router is a turntable, and you can bolt it in place

The router is a small linear layer. It scores every expert for the current token, turns the scores into probabilities with softmax, keeps the top k, and blends their outputs using those probabilities as weights.

logits  = router(token)          # one score per expert
probs   = softmax(logits)
w, idx  = topk(probs, k)         # keep the best k
out     = Σ w[i] · expert[idx[i]](token)  (+ shared expert)

Two different "weights". Parameters are the numbers stored inside an expert; they set its size and are fixed after training. Routing weights are the 0 to 1 shares the router hands out per token at runtime. Adding parameters makes a brick bigger. Raising routing weight makes the router lean on it more.

Adding a specialist expert, in three phases

Router and expert towers

# Phase 1: freeze everything except the new expert
for p in model.parameters(): p.requires_grad = False
for p in moe.experts[NEW].parameters(): p.requires_grad = True
out = model(batch, force_expert_idx=NEW)   # router bypassed

# Phase 2: lock the new expert, unlock the router
for p in moe.experts[NEW].parameters(): p.requires_grad = False
for p in moe.router.parameters(): p.requires_grad = True
out = model(batch)                          # normal top-k routing

The cold-start trap

The router needs a new output row for the new expert. If that row starts random, the router may never select it, and an expert that is never selected sends no gradient back to tell the router it is useful. Common fixes: initialize the new row with a small positive bias, or add a temporary auxiliary loss that nudges domain tokens toward it.

Load balancing

Left alone, routers collapse onto a few favorite experts. Training adds a balancing pressure (an auxiliary loss, or in DeepSeek's case a per-expert bias adjustment) so tokens spread across experts. This matters for your specialist too: push too hard on balance and medical tokens get scattered away from it.

Realistic note: on a 512-expert production model, bolting on a 513th expert means editing every MoE layer and the router. Most teams instead fine-tune with LoRA adapters on the existing experts and attention. The phased approach above is the right mental model and works well on smaller MoE models you train yourself.

4

Quantization: same studs, coarser plastic

Quantization keeps the count of numbers identical and stores each one in fewer bits. A 397B model is still 397B parameters at 4-bit. What changes is memory and precision.

How 4-bit actually works

Four bits can only express 16 distinct values. To make that usable, weights are split into small blocks (often 64 or 128 values). Each block stores its own scale, so a block of tiny numbers and a block of large numbers both get the full 16 levels of resolution.

Ten real weight values snapped to the nearest of sixteen evenly spaced levels Real weights in one block (BF16) Snapped to the nearest of 16 levels (4-bit) The block's scale stretches the 16 levels to fit its range. Small error per weight, about 3.5× less memory than BF16.
Each red dot is a real weight. It is stored as the nearest blue level plus one shared scale for the block.
real_value ≈ scale × q + bias (q is a 4-bit integer, 0 to 15)
MLX 4-bit, group size 64: 4 bits + (16-bit scale + 16-bit bias) / 64
= 4 + 0.5 = 4.5 effective bits per weight

That 0.5 bit of overhead is why a "4-bit" 671B model weighs about 378 GB rather than 335 GB.

What gets quantized and what doesn't

Expert weights are the bulk of the model and the main target. Routers, norms and often the shared expert and embeddings are commonly kept at higher precision, because small errors there affect every token. A bad router decision sends a token to the wrong specialist entirely.

FormatBytes per paramTypical useQuality impact
BF162Training, referenceBaseline
FP8 (block 128)1H100 serving, Qwen's official FP8 releaseNear lossless
8-bit integer~1.06Mac (MLX), llama.cpp Q8Near lossless
6-bit~0.81Mac sweet spot when it fitsSmall
4-bit~0.56Mac, consumer GPUsNoticeable on hard reasoning and long code, fine for most chat
3-bit and below~0.44Squeezing giants onto fewer machinesReal degradation, test carefully

Quantization speeds up generation too. Each new token requires reading the active weights from memory, so fewer bytes per weight means fewer bytes moved per token. Module 5 explains why that is the real speed limit.

5

Hardware: capacity decides fit, bandwidth decides speed

Generating a token means streaming every active weight through the chip once. So for one user, tokens per second is capped by memory bandwidth divided by bytes read per token. Raw compute matters more for reading long prompts and serving many users at once.

ceiling tok/s ≈ bandwidth ÷ (active params × bytes per param)
Qwen3.5 at 4-bit on one M5 Ultra: 1,200 GB/s ÷ (17B × 0.56) ≈ 125 tok/s ceiling
Real speeds land well below the ceiling. Use it to compare setups.
Mac Studio M5 Ultra (2026)NVIDIA H100 SXM
Memory per unitUp to 512 GB unified80 GB HBM3
Memory bandwidth1.2 TB/s3.35 TB/s
Link between unitsThunderbolt 5 with RDMA, up to 4 machines in a full meshNVLink 900 GB/s inside an 8-GPU node, InfiniBand between nodes
SoftwareMLX, exo, llama.cpp, LM StudiovLLM, SGLang, TensorRT-LLM
StrengthHuge memory per dollar, quiet, privateThroughput, batching many users, FP8, training
WeaknessLower bandwidth, slower prompt processing, young cluster tooling80 GB each means many GPUs for big models, hourly cost

Three ways to split a model across machines

Pipeline parallel: machine A holds layers 1 to 30, machine B holds 31 to 60. Simple, but machines wait on each other. Tensor parallel: every layer is sliced across all machines, which work together on each token. Fast, but chatty, so the link must be low latency. Expert parallel: different experts live on different machines and tokens are shipped to wherever their experts live. Native to MoE and the standard choice on H100 nodes.

Pipeline, tensor and expert parallelism compared Pipeline: split by layers A: layers 1 to 30B: layers 31 to 60 Tensor: split every layer A: half of every layerB: half of every layer Syncs on every layer, so it needs low latency Expert: split by experts E1 to 128E129 to 256E257 to 384E385 to 512 Tokens are shipped to wherever their experts live
Real deployments often combine these, for example tensor parallel inside a node and expert parallel across nodes.

Why RDMA over Thunderbolt mattered for Macs

Tensor parallel sends small messages between machines for every layer of every token. Over regular networking each hop took roughly 300 microseconds. macOS 26.2 added RDMA over Thunderbolt 5, which brought that under 50 microseconds and made tensor parallel across Mac Studios practical. Early reports had 4 × M3 Ultra running Qwen3 235B around 32 tok/s and Kimi K2 Thinking (1T) around 30 tok/s with exo. The new M5 Ultra has 50% more bandwidth than the M3 Ultra.

6

Size your build

Pick a model and precision. The calculator estimates weight memory, how many machines you need, and the bandwidth ceiling.

Assumptions: 90% of a Mac's unified memory usable for the model (macOS limits GPU memory by default; it can be raised), 90% of each H100 usable, plus 10% headroom for KV cache and runtime. H100 counts rounded up to whole 8-GPU nodes. Estimates, not guarantees.

7

Two reference builds for a 300B+ model

Recommended model: Qwen3.5-397B-A17B. It clears your 300B bar, uses the same 4,096 hidden width from this course, ships under Apache 2.0 for commercial use, and has an official FP8 release. Its hybrid attention keeps only 15 full-attention layers with 2 KV heads, so long-context memory stays small (roughly 30 KB per token at BF16, about 4 GB for 128K tokens).

Local cluster of up to four Mac Studios in a full Thunderbolt mesh compared with a cloud node of eight H100 GPUs linked by NVLink Local officeCloud node M5M5M5M5H100H100H100H100H100H100H100H100 NVLink 900 GB/s Up to 4 × M5 Ultra 512 GBThunderbolt 5 RDMA mesh 8 × H100 80 GB = 640 GBRented by the hour
Every Mac must be cabled directly to every other Mac. Inside an H100 node, all GPUs share NVLink.

Local: Mac Studio cluster

SetupWhat fitsNotes
1 × M5 Ultra 512 GBQwen3.5-397B at 4-bit (~223 GB) or 6-bit (~322 GB)Simplest. No cluster tooling needed.
2 × M5 Ultra 512 GBQwen3.5 at 8-bit, DeepSeek V3.1 / Mistral Large 3 at 6 to 8-bit, Kimi K2.6 at 4-bitTensor parallel via exo or MLX distributed over TB5 RDMA.
4 × M5 Ultra 512 GBKimi K2.6 at 8-bit (3 machines), DeepSeek V4-Pro at 4-bit (3 machines) or 6-bit (4, tight)Current RDMA limit is 4 machines, all cabled to each other.
# Single machine, MLX
pip install mlx-lm
mlx_lm.server --model mlx-community/<Qwen3.5-397B-A17B 4-bit repo>

# Cluster: enable RDMA on each Mac from recovery (macOS 26.2+),
# cable every Mac to every other Mac over Thunderbolt 5,
# then run exo on each machine; it discovers peers and shards the model.

Pricing: the M5 Ultra Mac Studio starts at $5,499 with 96 GB. The 512 GB configurations ship in late October 2026; check Apple's configurator for current pricing.

Cloud: H100 nodes

SetupWhat fitsNotes
1 node, 8 × H100 (640 GB)Qwen3.5-397B FP8 (~397 GB) with room for many concurrent usersBF16 (~794 GB) does not fit. Use FP8.
2 nodes, 16 × H100DeepSeek V3.1, Mistral Large 3, GLM-5.2 and (tightly) Kimi K2.6 at FP8; Qwen3.5 at BF16Needs InfiniBand between nodes.
4 nodes, 32 × H100DeepSeek V4-Pro at FP8H200 (141 GB) or B200 nodes are often cheaper per model at this size.
# One 8×H100 node, SGLang, tensor + expert parallel
python3 -m sglang.launch_server \
  --model-path Qwen/Qwen3.5-397B-A17B-FP8 \
  --tp 8 --ep 8 \
  --reasoning-parser qwen3 --tool-call-parser qwen3_coder

Cost: H100s rent for roughly $2 to $12 per GPU-hour in 2026, with specialist clouds near the low end. One 8-GPU node running all month at $2.50 per GPU-hour is about $14,600.

Choosing between them

Mac Studios win on private, always-on inference for a small team, with no hourly bill, and keep sensitive data (such as PHI) on hardware you control. H100s win on throughput for many users, prompt-heavy workloads, and any fine-tuning. A practical split: serve locally on Macs, rent H100s by the hour when you train an adapter or specialist expert. Training needs several times the inference memory because of gradients and optimizer state.

8

Why training needs far more memory than inference

Inference holds the model and the KV cache. Training holds the model plus everything needed to change it: a precise master copy, a gradient for every weight, the optimizer's memory, and the saved activations from the forward pass.

Item, per weightInferenceFull training (BF16 + AdamW)
Working weights2 bytes (BF16) or ~0.56 (4-bit)2 bytes
FP32 master copynone4 bytes
Gradientnone2 bytes
Adam state (m and v)none8 bytes
Total0.5 to 2 bytesabout 16 bytes, plus activations
Inference uses 2 bytes per weight; full training uses 16 bytes per weight Inference: 2 bytes per weight W Full training: 16 bytes + activations WmastergradAdam m + v Each 20 px = 1 byte. 16 bytes × 397B weights ≈ 6.4 TB
Same weights, eight times the memory once you train them with Adam.
Qwen3.5-397B inference at FP8: 397B × 1 byte ≈ 397 GB (one 8 × H100 node)
Qwen3.5-397B full training: 397B × 16 bytes ≈ 6.4 TB before activations (100+ H100s)

Activations are the hidden cost. During the forward pass each layer's input is saved so the backward pass can use it. That memory grows with batch size × sequence length × layers. Inference discards each layer's output as soon as the next layer reads it.

MoE makes full training heavier still. All 512 experts need gradients and optimizer state even though each token only touches 11. The fix is to train fewer weights: freeze most of the model (Module 3) or use LoRA (Module 10).

9

Gradients, the backward pass and the optimizer

Training repeats one loop: forward pass, measure the loss, backward pass, update. Step through it below on a model with a single weight.

Model: prediction = w × 2. Target is 3, so the right answer is w = 1.5. Start at w = 0.5.

Loss over training steps

Gradient

For every weight, the gradient answers: if I nudge this weight up slightly, does the loss go up or down, and by how much? A gradient of −8 means increasing the weight lowers the loss, strongly. A real model has one gradient per weight, so gradients cost as much memory as the weights.

Backward pass

Most weights sit many layers away from the loss. The backward pass starts at the loss and walks back through the layers in reverse, using the chain rule. Each layer works out its share of the blame and passes the rest back to the layer before it. To do that, each layer needs the exact input it saw on the forward pass. That is why activations must be stored.

Forward pass runs left to right through the layers to the loss; backward pass runs right to left computing gradients Forward: compute, and save each layer's input input L1 L2 … L60 loss Backward: chain rule, a gradient for every weight It needs the saved inputs, so training stores activations
Forward builds the prediction. Backward assigns blame to every weight.

Optimizer state

Plain SGD moves each weight by learning rate × gradient and remembers nothing. Adam keeps two running averages per weight, for the whole run:

m, momentum: the recent average direction. It smooths noisy batches so weights don't zigzag. v, variance: the recent average of the squared gradient. Weights with big gradients take smaller steps, weights with tiny gradients take bigger ones.

Toggle the two optimizers above. SGD's step size follows the gradient. Adam takes roughly even steps. That per-weight memory is two FP32 numbers, 8 bytes, the largest item in the 16-byte budget.

10

LoRA and QLoRA: snap on thin bricks instead of rebuilding

LoRA freezes the original matrix W and trains two skinny matrices, B and A, whose product is the change. The adapter's output is added to the frozen output.

Frozen 4,096 by 1,024 matrix W plus trainable B (4,096 by 16) times A (16 by 1,024)
output = W·x + (α / r) · B·A·x
W: 4,096 × 1,024 = 4.19M params, frozen
B: 4,096 × 16 + A: 16 × 1,024 = 82K params, trainable (about 2%)

Rank r (16 here) sets the adapter's capacity. Common values are 8 to 64. B starts at zero, so at step 0 the adapter adds nothing and the model behaves exactly like the original. Only A and B get gradients and Adam state, so the 16-byte cost applies to about 1 to 2% of weights.

It works because the changes fine-tuning makes are mostly low-rank. Teaching a model denial patterns or your output format doesn't require rewriting every weight independently.

After training

Either merge B·A into W (no speed cost) or keep the adapter as a separate file of tens to a few hundred MB. Separate adapters let one base model serve many jobs: a denials adapter, a credentialing adapter, a payer-contract adapter, swapped per request.

QLoRA

Load the frozen base at 4-bit (Module 4) and keep A and B at full precision. It is the cheapest way to fine-tune a large model, with a small quality cost from the 4-bit base.

ModelFull fine-tuneLoRA (BF16 base)QLoRA (4-bit base)
Qwen3.5-9B~150 GB, 2 to 4 H100s~25 to 35 GB, 1 GPU~10 to 16 GB, any 24 GB GPU or a Mac
Qwen3.5-35B-A3B~560 GB, 8+ H100s~80 to 100 GB, 1 to 2 H100s~25 to 35 GB, 1 H100 or a 64 GB+ Mac
Qwen3.5-397B-A17B~6.4 TB, 100+ H100s~850 GB+, 2 nodes~260 to 320 GB, 4 to 5 H100s

Rough figures for short sequences with gradient checkpointing. Long documents raise activation memory.

Other memory savers

Gradient checkpointing: store fewer activations and recompute them on the backward pass, about 30% more compute for much less memory. ZeRO or FSDP: split weights, gradients and optimizer state across GPUs instead of copying them to each one.

11

Worked example: a small fine-tuned denial triage model

A narrow, high-volume job is where a small fine-tuned model earns its keep. Here is one built for RCM.

Fine-tuning pipeline from resolved denials to deployed adapter with a correction loopResolved denials: 835 lines + notesDe-identify (Safe Harbor, 18 identifiers)Billers label cause, action, ownerQLoRA on Qwen3.5-9B, 1 to 3 GPU hoursEval gate: 300 held-out casesDeploy adapter, log every correction corrections feed back
The loop at the end is what keeps the model current: every human correction becomes future training data.
Spec
JobRead one denied claim line, return root cause, category, appealability, next action and owner, as strict JSON
Base modelQwen3.5-9B (dense, open weights; confirm license on the model card)
MethodQLoRA, rank 16, all linear layers, 2 to 3 epochs
Data5,000 to 20,000 de-identified denial lines your billers already resolved, plus 300 held back for testing
Training hardware1 rented H100 for 1 to 3 hours (roughly $5 to $40), or an M5 Max/Ultra Mac overnight
Serving4-bit, about 5 to 6 GB. Runs on any recent Mac at high speed, alongside the big model

One training example

{"messages": [
 {"role": "system", "content": "Classify the denial. Reply in JSON only."},
 {"role": "user", "content": "Payer: commercial PPO. CARC: CO-197. CPT: 72148. Billed: 1450.00. Paid: 0.00. Note: lumbar MRI, no auth on file."},
 {"role": "assistant", "content": "{\"root_cause\":\"missing_prior_auth\",\"category\":\"front_end\",\"appealable\":\"conditional\",\"next_action\":\"check payer retro-auth window, else adjust per contract\",\"owner\":\"auth_team\"}"}
]}

Train it

# On a Mac with MLX (data/ holds train.jsonl and valid.jsonl)
mlx_lm.lora --model <Qwen3.5-9B MLX repo> --train --data ./data \
  --fine-tune-type lora --batch-size 4 --iters 1500 \
  --adapter-path ./adapters/denials

# On a rented H100 with Hugging Face PEFT
LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
           target_modules="all-linear", task_type="CAUSAL_LM")
# load the base with BitsAndBytesConfig(load_in_4bit=True) for QLoRA

Benefits

Cheap: training costs tens of dollars, not thousands. Fast and private: it runs locally in milliseconds, so PHI never leaves your hardware. Consistent: it returns the same JSON shape every time, which makes downstream automation reliable. Your judgment, encoded: it learns how your team actually resolves denials, not generic advice. Stackable: one base, many adapters.

Drawbacks

Narrow: it is good at the one job and weak outside it. Weaker reasoning: a 9B model will miss multi-step logic a 397B model catches, so unusual cases need escalation. Goes stale: payer rules change and baked-in facts don't update, so retrain on a schedule. Copies your mistakes: inconsistent historical decisions become inconsistent model output. Data is the real cost: cleaning, de-identifying and labeling takes far longer than training.

The rule that saves the most money: fine-tune for behavior, retrieve for facts. Teach the model your format, workflow and judgment through fine-tuning. Feed it payer policies, fee schedules and code definitions at request time through retrieval (RAG), so updating a policy means editing a document, not retraining.

12

An RCM agent team on a sub-$100K budget

You do not need to pretrain anything. Pretraining is the million-dollar part. You need one strong general model for reasoning, a few small fine-tuned models for high-volume narrow jobs, good tools, and humans approving anything that touches money.

The team

RCM agent team: intake routes work to eligibility, coding, denial and credentialing agents, then appeals writer, QA reviewer and human approval, with shared memory and tools 835 / 837 / portal work Intake agent Eligibilitysmall modelCoding checklarge modelDenial triagesmall modelCredentialingsmall model Appeals writer QA reviewer Human approval then submit to payer Memorypayer quirkspast cases ToolsICD-10, CMSNPI, X12 small fine-tuned large model human
Dashed lines: every agent can read shared memory and call tools. Only humans release work to payers.
AgentJobModelTools
IntakeParse 835 remits and 837 claims, split work itemsCode, plus small model for messy notesX12 parser
EligibilityCheck coverage and benefits before and after serviceSmall fine-tunedClearinghouse or payer API
Coding checkFlag code mismatches and medical necessity gapsLarge model with RAGICD-10 lookup, CMS coverage (LCD/NCD)
Denial triageRoot cause and next action (the Module 11 model)Small fine-tunedDenial history
Appeals writerDraft appeal letters with cited policyLarge model with RAGPayer policy library, chart excerpts
CredentialingTrack enrollment status, flag expirationsSmall fine-tunedNPI Registry, CAQH, payer portals
QA reviewerCheck every output against rules before a human sees itLarge modelRule checklist

Large model: Qwen3.5-397B on one 512 GB Mac Studio at 4-bit, or a frontier API for the hardest cases. Small models: Qwen3.5-9B adapters on the same machine. A human approves submissions, write-offs and appeals.

Year-one budget (estimates)

ItemRangeNotes
1 × Mac Studio M5 Ultra, 512 GB~$12K to $18KServes the large model and all adapters. 512 GB price not yet listed; confirm with Apple
Cloud GPU bursts for training$500 to $5KRented H100s by the hour, only during training runs
Frontier API for hard cases$2K to $15KDepends on volume. Requires a signed BAA before any PHI is sent
Engineering$0 to $50KYour time, or a contractor for the agent framework and integrations
Data labelingstaff hoursBillers reviewing and correcting examples. The biggest real cost
Total~$15K to $90K + staff timeAbout 2 to 9% of a $1M budget

Sequence

Weeks 1 to 4: pick three tasks, pull and de-identify historical data, build a 300-example graded test set. Months 2 to 3: run agents on the large model with RAG and no fine-tuning, and measure against the test set. Months 3 to 4: fine-tune small models only for tasks that are high-volume and where the large model is slow, costly or inconsistent. Ongoing: log every human correction, add it to the training data, retrain adapters quarterly.

Compliance checklist: keep PHI on hardware you control or with vendors under a BAA. De-identify training data (HIPAA Safe Harbor's 18 identifiers). CPT descriptors are licensed by the AMA, so check your license before training on code descriptions. Keep a human approval step on anything submitted to a payer.

13

Persistent memory: how agents remember between sessions

A model's weights are frozen once it is deployed, and its context window empties when a session ends. Everything an agent "remembers" next week has to be written somewhere and read back in. That store is persistent memory, and in 2026 it is a core design decision for any agent team.

Five layers of agent memory from fastest changing to slowest: working, episodic, semantic, procedural, weightsWorking (context window)this conversation, then goneevery turnEpisodicwhat happened: case historyevery caseSemanticwhat is true: payer factswhen facts changeProceduralhow we do it: playbooksafter reviewWeights (LoRA)baked-in behaviorquarterly retrain updates
Faster layers are cheap to change and easy to get wrong. Slower layers are durable and harder to change.

The four kinds

Working memory is the context window: the prompt, documents and conversation the model sees right now. Episodic memory is a log of what happened: this claim, this payer, this appeal, this outcome. Semantic memory is distilled facts believed to be true now: "Payer X wants modifier 25 documentation attached." Procedural memory is know-how: playbooks, prompts and skills that tell an agent how to do a job. LoRA weights are the slowest layer, behavior baked into the model itself.

How it works in practice

Memories are stored in a database (often Postgres with a vector index, which Supabase provides), retrieved by similarity or by structured lookup, and inserted into the context window before the model answers. A background process consolidates: it turns many episodes into one semantic fact ("the last 12 appeals to Payer X won when we attached the op note"), and it prunes or corrects facts that went stale.

ApproachHow it worksExample tools
Agent-managedThe agent decides what to write and edit in its own memory blocksLetta (formerly MemGPT)
Pipeline-managedA separate process extracts facts from each conversation and stores themMem0
Time-awareFacts carry valid-from and valid-to dates so old rules are not mistaken for current onesZep
Build your ownTables for cases, facts and playbooks, plus a vector indexSupabase or Postgres with pgvector

Memory for an RCM team

Episodic: every work item with its payer, codes, action taken and outcome. Semantic: payer rules by payer and state, client preferences, contract terms, each with a source and a date. Procedural: one playbook per task, for example "denial CO-197 at Payer X", that agents follow and propose edits to. Weights: the quarterly LoRA adapters from Module 11.

Three rules that prevent most memory failures. Store a source and a date with every fact, so stale rules can be found and expired. Never let an agent rewrite a playbook without a test run and a human sign-off. Treat the memory store as PHI: same access controls, BAA and retention policy as your claims data. A memory store is also an attack surface; text from a payer portal or email should never be saved as an instruction.

14

Reinforcement learning: training on outcomes, not examples

Supervised fine-tuning (Module 11) shows the model correct answers and asks it to imitate them. Reinforcement learning lets the model try, scores the attempt, and makes higher-scoring behavior more likely. It is how models learned to reason step by step, and it can now run on a single GPU for small models.

StageSignalWhat it teaches
PretrainingPredict the next token on trillions of tokensLanguage and world knowledge. The million-dollar part; you skip it.
Supervised fine-tuning (SFT)Your examples of correct outputFormat, workflow, your team's style
Preference tuning (RLHF, DPO)A says which of two answers is betterTone, helpfulness, judgment calls
RL with verifiable rewards (RLVR)A program checks whether the answer is rightGetting the right answer more often, multi-step reasoning

GRPO, the default RL method in 2026

GRPO: one prompt, four sampled answers, a verifier scores them, and each answer's advantage is its reward minus the group average Prompt: one denial line Answer Areward 1.0adv +0.5Answer Breward 0.2adv −0.3Answer Creward 0.8adv +0.3Answer Dreward 0.0adv −0.5 Verifier: valid JSON? right cause? matches outcome? Group average is 0.5. Above average gets reinforced, below average gets discouraged. No critic model needed.
Group Relative Policy Optimization compares answers to each other, so it needs no separate critic model. That saves a lot of memory.

For each prompt the model generates several answers. A reward function scores each one. Each answer's advantage is its score minus the group average. Training then raises the probability of above-average answers and lowers the rest. A penalty (the KL term) keeps the model close to where it started so it does not drift or forget its general skills. With LoRA, the frozen base model doubles as that reference, which keeps memory low. Libraries such as Unsloth and Hugging Face TRL run GRPO on a single 24 GB GPU for small models.

Verifiable rewards for RCM

RL only works if you can score answers automatically. RCM has more of these than most fields: is the output valid JSON; does the predicted denial category match how the claim was finally resolved; does the CARC-to-action mapping match your rules table; did the appeal the model recommended actually get overturned. That last one is a delayed reward, so you collect outcomes and train in batches.

Reward hacking is the main risk. Reward the number of appeals filed and the model learns to call everything appealable. Reward outcomes (dollars recovered, overturn rate, rework avoided), cap the reward for doing nothing risky, and review samples by hand every training round. The practical order: SFT first so the model knows the format, then GRPO to sharpen accuracy on cases you can verify.

15

Recursive self-improvement: where it stands and where it is going

Recursive self-improvement (RSI) means AI systems doing the work of building better AI systems, which then do that work even better. The full version, an AI that chooses its own research goals and builds its successor without humans, has not been demonstrated. Large parts of the loop are now automated.

Three stages of AI doing AI research: done, happening now, not yet Who runs the AI-building loop 2023 to 2025AI suggests code,humans run it 2026: nowAI writes mostcode and runsexperimentsHumans stillchoose goals Not yetAI picks goalsand builds itsown successorThis is fullrecursive self-improvement
Based on public reporting from Anthropic, OpenAI and independent researchers as of September 2026.

Current state, September 2026

Anthropic reports that more than 80% of code merged into its codebase was written by Claude as of May 2026, up from low single digits before February 2025, and that its engineers merge about 8 times as much code per day as in 2024. On a fixed speed-optimization task, Claude went from about a 3× speedup in May 2025 to about 52× in April 2026; a skilled human reaches about 4× in four to eight hours. In an open-ended AI safety research problem, Claude agents closed 97% of a performance gap that two human researchers closed 23% of in a week, using about $18,000 of compute, though humans chose the problem and the scoring.

OpenAI said in February 2026 that early versions of GPT-5.3-Codex helped debug its own training and evaluation. Narrower self-improving systems, such as DeepMind's AlphaEvolve and the Darwin Gödel Machine, improve code or prompts against a fixed score they do not control. Both Anthropic and OpenAI say fully autonomous RSI is not happening yet. The remaining human role is choosing which problems matter and judging which results to trust.

Expected trajectory

METR measures the length of tasks AI can complete on its own; that length has been doubling about every four months. Anthropic projects that tasks taking a skilled person days could come into range in 2026, and tasks taking weeks in 2027. It describes three possible futures:

ScenarioWhat happensAnthropic's view
Stall and diffuseProgress flattens, but today's models still spread through the economyPossible, not likely
Compounding efficiencyAI does most of the work, humans set direction and verify. A 100-person firm does the work of thousandsThe likely path
Full RSIAI designs and trains its successors; humans move to oversightPlausible, timing and safety uncertain

Even in the faster scenarios, bottlenecks move rather than disappear. Anthropic already finds human code review is its new constraint. For a business, the equivalent is approval, verification and payer response times, not drafting speed.

What this means for your RCM operation

You can run a bounded, safe version of the same loop today. Agents propose changes to their own playbooks and prompts. Each change is scored against a fixed test set you control (the 300 graded cases from Module 12). Only changes that score higher, and that a human approves, are kept. Corrections feed the next LoRA or GRPO round. That is recursive improvement of your agent team, with the scoring and the final say kept in human hands.

Not to be confused with recursive language models, a separate long-context technique in which a model breaks a huge document into pieces and calls itself on each piece. Useful for 300-page payer manuals, but unrelated to self-improvement.

16

Quiz

Tap an answer to see the explanation. Your progress is saved in this browser.

0 of 35 correct