A token is a row of 4,096 numbers
Text is chopped into tokens: whole words or pieces of words. "Credentialing" might become three tokens. Each token gets an ID from the vocabulary, and that ID is looked up in an embedding table, which returns a list of decimal numbers. From then on, that list is the token.
The length of that list is the model's hidden dimension, written d_model. Every token has the same width, at every layer, whether it is "the", a comma or "cardiology". A long word does not get a wider vector. It gets split into more tokens.
| Model | Hidden width (d_model) |
|---|---|
| Mixtral 8x7B, Mistral 7B | 4,096 |
| Qwen3.5-397B-A17B (our main example) | 4,096 |
| Mixtral 8x22B | 6,144 |
| DeepSeek V3, Kimi K2 | 7,168 |
The 4,096 values are numbers, not bits. Each number is stored using some number of bits: 16 in BF16, 8 in FP8, about 4 when quantized. One token at BF16 is 4,096 × 16 bits = 8 KB. Module 4 builds on this.
An expert is a brick with a fixed stud pattern
Every expert is a small feed-forward network. Its input and output must be exactly d_model wide so it snaps onto the rest of the model. Inside, it expands the token to a wider working space, filters it, and compresses it back. That inner width is d_ff, the intermediate size.
token (4,096) ─▶ gate proj ─┐ token (4,096) ─▶ up proj ─┴▶ SwiGLU filter (d_ff) ─▶ down proj ─▶ out (4,096)
Three matrices, so parameters per expert per layer = 3 × d_model × d_ff.
Coarse experts versus fine-grained experts
Mixtral uses a few big bricks. Newer models use many small ones, which lets the router mix specialists more precisely.
| Mixtral 8x7B | Qwen3.5-397B | |
|---|---|---|
| Experts per layer | 8 | 512 |
| Active per token | 2 | 10 routed + 1 shared |
| Expert inner width (d_ff) | 14,336 | 1,024 |
| Params per expert per layer | ~176M | ~12.6M |
| Layers | 32 | 60 |
Check the math yourself
512 experts × 12.58M = 6.44B per layer
6.44B × 60 layers ≈ 386B in experts
+ attention, embeddings, vision encoder ≈ 397B total
Active per token: 11 experts × 12.58M × 60 ≈ 8.3B
+ attention and embeddings ≈ 17B active
That is the whole trick of MoE. You store 397B parameters, but each token only touches about 17B of them. Memory scales with the total. Compute per token scales with the active count.
A shared expert is one brick every token always passes through. It holds common knowledge (grammar, general facts) so the routed experts can specialize.
The router is a turntable, and you can bolt it in place
The router is a small linear layer. It scores every expert for the current token, turns the scores into probabilities with softmax, keeps the top k, and blends their outputs using those probabilities as weights.
logits = router(token) # one score per expert probs = softmax(logits) w, idx = topk(probs, k) # keep the best k out = Σ w[i] · expert[idx[i]](token) (+ shared expert)
Two different "weights". Parameters are the numbers stored inside an expert; they set its size and are fixed after training. Routing weights are the 0 to 1 shares the router hands out per token at runtime. Adding parameters makes a brick bigger. Raising routing weight makes the router lean on it more.
Adding a specialist expert, in three phases
# Phase 1: freeze everything except the new expert for p in model.parameters(): p.requires_grad = False for p in moe.experts[NEW].parameters(): p.requires_grad = True out = model(batch, force_expert_idx=NEW) # router bypassed # Phase 2: lock the new expert, unlock the router for p in moe.experts[NEW].parameters(): p.requires_grad = False for p in moe.router.parameters(): p.requires_grad = True out = model(batch) # normal top-k routing
The cold-start trap
The router needs a new output row for the new expert. If that row starts random, the router may never select it, and an expert that is never selected sends no gradient back to tell the router it is useful. Common fixes: initialize the new row with a small positive bias, or add a temporary auxiliary loss that nudges domain tokens toward it.
Load balancing
Left alone, routers collapse onto a few favorite experts. Training adds a balancing pressure (an auxiliary loss, or in DeepSeek's case a per-expert bias adjustment) so tokens spread across experts. This matters for your specialist too: push too hard on balance and medical tokens get scattered away from it.
Realistic note: on a 512-expert production model, bolting on a 513th expert means editing every MoE layer and the router. Most teams instead fine-tune with LoRA adapters on the existing experts and attention. The phased approach above is the right mental model and works well on smaller MoE models you train yourself.
Quantization: same studs, coarser plastic
Quantization keeps the count of numbers identical and stores each one in fewer bits. A 397B model is still 397B parameters at 4-bit. What changes is memory and precision.
How 4-bit actually works
Four bits can only express 16 distinct values. To make that usable, weights are split into small blocks (often 64 or 128 values). Each block stores its own scale, so a block of tiny numbers and a block of large numbers both get the full 16 levels of resolution.
MLX 4-bit, group size 64: 4 bits + (16-bit scale + 16-bit bias) / 64
= 4 + 0.5 = 4.5 effective bits per weight
That 0.5 bit of overhead is why a "4-bit" 671B model weighs about 378 GB rather than 335 GB.
What gets quantized and what doesn't
Expert weights are the bulk of the model and the main target. Routers, norms and often the shared expert and embeddings are commonly kept at higher precision, because small errors there affect every token. A bad router decision sends a token to the wrong specialist entirely.
| Format | Bytes per param | Typical use | Quality impact |
|---|---|---|---|
| BF16 | 2 | Training, reference | Baseline |
| FP8 (block 128) | 1 | H100 serving, Qwen's official FP8 release | Near lossless |
| 8-bit integer | ~1.06 | Mac (MLX), llama.cpp Q8 | Near lossless |
| 6-bit | ~0.81 | Mac sweet spot when it fits | Small |
| 4-bit | ~0.56 | Mac, consumer GPUs | Noticeable on hard reasoning and long code, fine for most chat |
| 3-bit and below | ~0.44 | Squeezing giants onto fewer machines | Real degradation, test carefully |
Quantization speeds up generation too. Each new token requires reading the active weights from memory, so fewer bytes per weight means fewer bytes moved per token. Module 5 explains why that is the real speed limit.
Hardware: capacity decides fit, bandwidth decides speed
Generating a token means streaming every active weight through the chip once. So for one user, tokens per second is capped by memory bandwidth divided by bytes read per token. Raw compute matters more for reading long prompts and serving many users at once.
Qwen3.5 at 4-bit on one M5 Ultra: 1,200 GB/s ÷ (17B × 0.56) ≈ 125 tok/s ceiling
Real speeds land well below the ceiling. Use it to compare setups.
| Mac Studio M5 Ultra (2026) | NVIDIA H100 SXM | |
|---|---|---|
| Memory per unit | Up to 512 GB unified | 80 GB HBM3 |
| Memory bandwidth | 1.2 TB/s | 3.35 TB/s |
| Link between units | Thunderbolt 5 with RDMA, up to 4 machines in a full mesh | NVLink 900 GB/s inside an 8-GPU node, InfiniBand between nodes |
| Software | MLX, exo, llama.cpp, LM Studio | vLLM, SGLang, TensorRT-LLM |
| Strength | Huge memory per dollar, quiet, private | Throughput, batching many users, FP8, training |
| Weakness | Lower bandwidth, slower prompt processing, young cluster tooling | 80 GB each means many GPUs for big models, hourly cost |
Three ways to split a model across machines
Pipeline parallel: machine A holds layers 1 to 30, machine B holds 31 to 60. Simple, but machines wait on each other. Tensor parallel: every layer is sliced across all machines, which work together on each token. Fast, but chatty, so the link must be low latency. Expert parallel: different experts live on different machines and tokens are shipped to wherever their experts live. Native to MoE and the standard choice on H100 nodes.
Why RDMA over Thunderbolt mattered for Macs
Tensor parallel sends small messages between machines for every layer of every token. Over regular networking each hop took roughly 300 microseconds. macOS 26.2 added RDMA over Thunderbolt 5, which brought that under 50 microseconds and made tensor parallel across Mac Studios practical. Early reports had 4 × M3 Ultra running Qwen3 235B around 32 tok/s and Kimi K2 Thinking (1T) around 30 tok/s with exo. The new M5 Ultra has 50% more bandwidth than the M3 Ultra.
Size your build
Pick a model and precision. The calculator estimates weight memory, how many machines you need, and the bandwidth ceiling.
Assumptions: 90% of a Mac's unified memory usable for the model (macOS limits GPU memory by default; it can be raised), 90% of each H100 usable, plus 10% headroom for KV cache and runtime. H100 counts rounded up to whole 8-GPU nodes. Estimates, not guarantees.
Two reference builds for a 300B+ model
Recommended model: Qwen3.5-397B-A17B. It clears your 300B bar, uses the same 4,096 hidden width from this course, ships under Apache 2.0 for commercial use, and has an official FP8 release. Its hybrid attention keeps only 15 full-attention layers with 2 KV heads, so long-context memory stays small (roughly 30 KB per token at BF16, about 4 GB for 128K tokens).
Local: Mac Studio cluster
| Setup | What fits | Notes |
|---|---|---|
| 1 × M5 Ultra 512 GB | Qwen3.5-397B at 4-bit (~223 GB) or 6-bit (~322 GB) | Simplest. No cluster tooling needed. |
| 2 × M5 Ultra 512 GB | Qwen3.5 at 8-bit, DeepSeek V3.1 / Mistral Large 3 at 6 to 8-bit, Kimi K2.6 at 4-bit | Tensor parallel via exo or MLX distributed over TB5 RDMA. |
| 4 × M5 Ultra 512 GB | Kimi K2.6 at 8-bit (3 machines), DeepSeek V4-Pro at 4-bit (3 machines) or 6-bit (4, tight) | Current RDMA limit is 4 machines, all cabled to each other. |
# Single machine, MLX pip install mlx-lm mlx_lm.server --model mlx-community/<Qwen3.5-397B-A17B 4-bit repo> # Cluster: enable RDMA on each Mac from recovery (macOS 26.2+), # cable every Mac to every other Mac over Thunderbolt 5, # then run exo on each machine; it discovers peers and shards the model.
Pricing: the M5 Ultra Mac Studio starts at $5,499 with 96 GB. The 512 GB configurations ship in late October 2026; check Apple's configurator for current pricing.
Cloud: H100 nodes
| Setup | What fits | Notes |
|---|---|---|
| 1 node, 8 × H100 (640 GB) | Qwen3.5-397B FP8 (~397 GB) with room for many concurrent users | BF16 (~794 GB) does not fit. Use FP8. |
| 2 nodes, 16 × H100 | DeepSeek V3.1, Mistral Large 3, GLM-5.2 and (tightly) Kimi K2.6 at FP8; Qwen3.5 at BF16 | Needs InfiniBand between nodes. |
| 4 nodes, 32 × H100 | DeepSeek V4-Pro at FP8 | H200 (141 GB) or B200 nodes are often cheaper per model at this size. |
# One 8×H100 node, SGLang, tensor + expert parallel python3 -m sglang.launch_server \ --model-path Qwen/Qwen3.5-397B-A17B-FP8 \ --tp 8 --ep 8 \ --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Cost: H100s rent for roughly $2 to $12 per GPU-hour in 2026, with specialist clouds near the low end. One 8-GPU node running all month at $2.50 per GPU-hour is about $14,600.
Choosing between them
Mac Studios win on private, always-on inference for a small team, with no hourly bill, and keep sensitive data (such as PHI) on hardware you control. H100s win on throughput for many users, prompt-heavy workloads, and any fine-tuning. A practical split: serve locally on Macs, rent H100s by the hour when you train an adapter or specialist expert. Training needs several times the inference memory because of gradients and optimizer state.
Why training needs far more memory than inference
Inference holds the model and the KV cache. Training holds the model plus everything needed to change it: a precise master copy, a gradient for every weight, the optimizer's memory, and the saved activations from the forward pass.
| Item, per weight | Inference | Full training (BF16 + AdamW) |
|---|---|---|
| Working weights | 2 bytes (BF16) or ~0.56 (4-bit) | 2 bytes |
| FP32 master copy | none | 4 bytes |
| Gradient | none | 2 bytes |
| Adam state (m and v) | none | 8 bytes |
| Total | 0.5 to 2 bytes | about 16 bytes, plus activations |
Qwen3.5-397B full training: 397B × 16 bytes ≈ 6.4 TB before activations (100+ H100s)
Activations are the hidden cost. During the forward pass each layer's input is saved so the backward pass can use it. That memory grows with batch size × sequence length × layers. Inference discards each layer's output as soon as the next layer reads it.
MoE makes full training heavier still. All 512 experts need gradients and optimizer state even though each token only touches 11. The fix is to train fewer weights: freeze most of the model (Module 3) or use LoRA (Module 10).
Gradients, the backward pass and the optimizer
Training repeats one loop: forward pass, measure the loss, backward pass, update. Step through it below on a model with a single weight.
Model: prediction = w × 2. Target is 3, so the right answer is w = 1.5. Start at w = 0.5.
Gradient
For every weight, the gradient answers: if I nudge this weight up slightly, does the loss go up or down, and by how much? A gradient of −8 means increasing the weight lowers the loss, strongly. A real model has one gradient per weight, so gradients cost as much memory as the weights.
Backward pass
Most weights sit many layers away from the loss. The backward pass starts at the loss and walks back through the layers in reverse, using the chain rule. Each layer works out its share of the blame and passes the rest back to the layer before it. To do that, each layer needs the exact input it saw on the forward pass. That is why activations must be stored.
Optimizer state
Plain SGD moves each weight by learning rate × gradient and remembers nothing. Adam keeps two running averages per weight, for the whole run:
m, momentum: the recent average direction. It smooths noisy batches so weights don't zigzag. v, variance: the recent average of the squared gradient. Weights with big gradients take smaller steps, weights with tiny gradients take bigger ones.
Toggle the two optimizers above. SGD's step size follows the gradient. Adam takes roughly even steps. That per-weight memory is two FP32 numbers, 8 bytes, the largest item in the 16-byte budget.
LoRA and QLoRA: snap on thin bricks instead of rebuilding
LoRA freezes the original matrix W and trains two skinny matrices, B and A, whose product is the change. The adapter's output is added to the frozen output.
W: 4,096 × 1,024 = 4.19M params, frozen
B: 4,096 × 16 + A: 16 × 1,024 = 82K params, trainable (about 2%)
Rank r (16 here) sets the adapter's capacity. Common values are 8 to 64. B starts at zero, so at step 0 the adapter adds nothing and the model behaves exactly like the original. Only A and B get gradients and Adam state, so the 16-byte cost applies to about 1 to 2% of weights.
It works because the changes fine-tuning makes are mostly low-rank. Teaching a model denial patterns or your output format doesn't require rewriting every weight independently.
After training
Either merge B·A into W (no speed cost) or keep the adapter as a separate file of tens to a few hundred MB. Separate adapters let one base model serve many jobs: a denials adapter, a credentialing adapter, a payer-contract adapter, swapped per request.
QLoRA
Load the frozen base at 4-bit (Module 4) and keep A and B at full precision. It is the cheapest way to fine-tune a large model, with a small quality cost from the 4-bit base.
| Model | Full fine-tune | LoRA (BF16 base) | QLoRA (4-bit base) |
|---|---|---|---|
| Qwen3.5-9B | ~150 GB, 2 to 4 H100s | ~25 to 35 GB, 1 GPU | ~10 to 16 GB, any 24 GB GPU or a Mac |
| Qwen3.5-35B-A3B | ~560 GB, 8+ H100s | ~80 to 100 GB, 1 to 2 H100s | ~25 to 35 GB, 1 H100 or a 64 GB+ Mac |
| Qwen3.5-397B-A17B | ~6.4 TB, 100+ H100s | ~850 GB+, 2 nodes | ~260 to 320 GB, 4 to 5 H100s |
Rough figures for short sequences with gradient checkpointing. Long documents raise activation memory.
Other memory savers
Gradient checkpointing: store fewer activations and recompute them on the backward pass, about 30% more compute for much less memory. ZeRO or FSDP: split weights, gradients and optimizer state across GPUs instead of copying them to each one.
Worked example: a small fine-tuned denial triage model
A narrow, high-volume job is where a small fine-tuned model earns its keep. Here is one built for RCM.
| Spec | |
|---|---|
| Job | Read one denied claim line, return root cause, category, appealability, next action and owner, as strict JSON |
| Base model | Qwen3.5-9B (dense, open weights; confirm license on the model card) |
| Method | QLoRA, rank 16, all linear layers, 2 to 3 epochs |
| Data | 5,000 to 20,000 de-identified denial lines your billers already resolved, plus 300 held back for testing |
| Training hardware | 1 rented H100 for 1 to 3 hours (roughly $5 to $40), or an M5 Max/Ultra Mac overnight |
| Serving | 4-bit, about 5 to 6 GB. Runs on any recent Mac at high speed, alongside the big model |
One training example
{"messages": [
{"role": "system", "content": "Classify the denial. Reply in JSON only."},
{"role": "user", "content": "Payer: commercial PPO. CARC: CO-197. CPT: 72148. Billed: 1450.00. Paid: 0.00. Note: lumbar MRI, no auth on file."},
{"role": "assistant", "content": "{\"root_cause\":\"missing_prior_auth\",\"category\":\"front_end\",\"appealable\":\"conditional\",\"next_action\":\"check payer retro-auth window, else adjust per contract\",\"owner\":\"auth_team\"}"}
]}
Train it
# On a Mac with MLX (data/ holds train.jsonl and valid.jsonl)
mlx_lm.lora --model <Qwen3.5-9B MLX repo> --train --data ./data \
--fine-tune-type lora --batch-size 4 --iters 1500 \
--adapter-path ./adapters/denials
# On a rented H100 with Hugging Face PEFT
LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", task_type="CAUSAL_LM")
# load the base with BitsAndBytesConfig(load_in_4bit=True) for QLoRA
Benefits
Cheap: training costs tens of dollars, not thousands. Fast and private: it runs locally in milliseconds, so PHI never leaves your hardware. Consistent: it returns the same JSON shape every time, which makes downstream automation reliable. Your judgment, encoded: it learns how your team actually resolves denials, not generic advice. Stackable: one base, many adapters.
Drawbacks
Narrow: it is good at the one job and weak outside it. Weaker reasoning: a 9B model will miss multi-step logic a 397B model catches, so unusual cases need escalation. Goes stale: payer rules change and baked-in facts don't update, so retrain on a schedule. Copies your mistakes: inconsistent historical decisions become inconsistent model output. Data is the real cost: cleaning, de-identifying and labeling takes far longer than training.
The rule that saves the most money: fine-tune for behavior, retrieve for facts. Teach the model your format, workflow and judgment through fine-tuning. Feed it payer policies, fee schedules and code definitions at request time through retrieval (RAG), so updating a policy means editing a document, not retraining.
An RCM agent team on a sub-$100K budget
You do not need to pretrain anything. Pretraining is the million-dollar part. You need one strong general model for reasoning, a few small fine-tuned models for high-volume narrow jobs, good tools, and humans approving anything that touches money.
The team
| Agent | Job | Model | Tools |
|---|---|---|---|
| Intake | Parse 835 remits and 837 claims, split work items | Code, plus small model for messy notes | X12 parser |
| Eligibility | Check coverage and benefits before and after service | Small fine-tuned | Clearinghouse or payer API |
| Coding check | Flag code mismatches and medical necessity gaps | Large model with RAG | ICD-10 lookup, CMS coverage (LCD/NCD) |
| Denial triage | Root cause and next action (the Module 11 model) | Small fine-tuned | Denial history |
| Appeals writer | Draft appeal letters with cited policy | Large model with RAG | Payer policy library, chart excerpts |
| Credentialing | Track enrollment status, flag expirations | Small fine-tuned | NPI Registry, CAQH, payer portals |
| QA reviewer | Check every output against rules before a human sees it | Large model | Rule checklist |
Large model: Qwen3.5-397B on one 512 GB Mac Studio at 4-bit, or a frontier API for the hardest cases. Small models: Qwen3.5-9B adapters on the same machine. A human approves submissions, write-offs and appeals.
Year-one budget (estimates)
| Item | Range | Notes |
|---|---|---|
| 1 × Mac Studio M5 Ultra, 512 GB | ~$12K to $18K | Serves the large model and all adapters. 512 GB price not yet listed; confirm with Apple |
| Cloud GPU bursts for training | $500 to $5K | Rented H100s by the hour, only during training runs |
| Frontier API for hard cases | $2K to $15K | Depends on volume. Requires a signed BAA before any PHI is sent |
| Engineering | $0 to $50K | Your time, or a contractor for the agent framework and integrations |
| Data labeling | staff hours | Billers reviewing and correcting examples. The biggest real cost |
| Total | ~$15K to $90K + staff time | About 2 to 9% of a $1M budget |
Sequence
Weeks 1 to 4: pick three tasks, pull and de-identify historical data, build a 300-example graded test set. Months 2 to 3: run agents on the large model with RAG and no fine-tuning, and measure against the test set. Months 3 to 4: fine-tune small models only for tasks that are high-volume and where the large model is slow, costly or inconsistent. Ongoing: log every human correction, add it to the training data, retrain adapters quarterly.
Compliance checklist: keep PHI on hardware you control or with vendors under a BAA. De-identify training data (HIPAA Safe Harbor's 18 identifiers). CPT descriptors are licensed by the AMA, so check your license before training on code descriptions. Keep a human approval step on anything submitted to a payer.
Persistent memory: how agents remember between sessions
A model's weights are frozen once it is deployed, and its context window empties when a session ends. Everything an agent "remembers" next week has to be written somewhere and read back in. That store is persistent memory, and in 2026 it is a core design decision for any agent team.
The four kinds
Working memory is the context window: the prompt, documents and conversation the model sees right now. Episodic memory is a log of what happened: this claim, this payer, this appeal, this outcome. Semantic memory is distilled facts believed to be true now: "Payer X wants modifier 25 documentation attached." Procedural memory is know-how: playbooks, prompts and skills that tell an agent how to do a job. LoRA weights are the slowest layer, behavior baked into the model itself.
How it works in practice
Memories are stored in a database (often Postgres with a vector index, which Supabase provides), retrieved by similarity or by structured lookup, and inserted into the context window before the model answers. A background process consolidates: it turns many episodes into one semantic fact ("the last 12 appeals to Payer X won when we attached the op note"), and it prunes or corrects facts that went stale.
| Approach | How it works | Example tools |
|---|---|---|
| Agent-managed | The agent decides what to write and edit in its own memory blocks | Letta (formerly MemGPT) |
| Pipeline-managed | A separate process extracts facts from each conversation and stores them | Mem0 |
| Time-aware | Facts carry valid-from and valid-to dates so old rules are not mistaken for current ones | Zep |
| Build your own | Tables for cases, facts and playbooks, plus a vector index | Supabase or Postgres with pgvector |
Memory for an RCM team
Episodic: every work item with its payer, codes, action taken and outcome. Semantic: payer rules by payer and state, client preferences, contract terms, each with a source and a date. Procedural: one playbook per task, for example "denial CO-197 at Payer X", that agents follow and propose edits to. Weights: the quarterly LoRA adapters from Module 11.
Three rules that prevent most memory failures. Store a source and a date with every fact, so stale rules can be found and expired. Never let an agent rewrite a playbook without a test run and a human sign-off. Treat the memory store as PHI: same access controls, BAA and retention policy as your claims data. A memory store is also an attack surface; text from a payer portal or email should never be saved as an instruction.
Reinforcement learning: training on outcomes, not examples
Supervised fine-tuning (Module 11) shows the model correct answers and asks it to imitate them. Reinforcement learning lets the model try, scores the attempt, and makes higher-scoring behavior more likely. It is how models learned to reason step by step, and it can now run on a single GPU for small models.
| Stage | Signal | What it teaches |
|---|---|---|
| Pretraining | Predict the next token on trillions of tokens | Language and world knowledge. The million-dollar part; you skip it. |
| Supervised fine-tuning (SFT) | Your examples of correct output | Format, workflow, your team's style |
| Preference tuning (RLHF, DPO) | A says which of two answers is better | Tone, helpfulness, judgment calls |
| RL with verifiable rewards (RLVR) | A program checks whether the answer is right | Getting the right answer more often, multi-step reasoning |
GRPO, the default RL method in 2026
For each prompt the model generates several answers. A reward function scores each one. Each answer's advantage is its score minus the group average. Training then raises the probability of above-average answers and lowers the rest. A penalty (the KL term) keeps the model close to where it started so it does not drift or forget its general skills. With LoRA, the frozen base model doubles as that reference, which keeps memory low. Libraries such as Unsloth and Hugging Face TRL run GRPO on a single 24 GB GPU for small models.
Verifiable rewards for RCM
RL only works if you can score answers automatically. RCM has more of these than most fields: is the output valid JSON; does the predicted denial category match how the claim was finally resolved; does the CARC-to-action mapping match your rules table; did the appeal the model recommended actually get overturned. That last one is a delayed reward, so you collect outcomes and train in batches.
Reward hacking is the main risk. Reward the number of appeals filed and the model learns to call everything appealable. Reward outcomes (dollars recovered, overturn rate, rework avoided), cap the reward for doing nothing risky, and review samples by hand every training round. The practical order: SFT first so the model knows the format, then GRPO to sharpen accuracy on cases you can verify.
Recursive self-improvement: where it stands and where it is going
Recursive self-improvement (RSI) means AI systems doing the work of building better AI systems, which then do that work even better. The full version, an AI that chooses its own research goals and builds its successor without humans, has not been demonstrated. Large parts of the loop are now automated.
Current state, September 2026
Anthropic reports that more than 80% of code merged into its codebase was written by Claude as of May 2026, up from low single digits before February 2025, and that its engineers merge about 8 times as much code per day as in 2024. On a fixed speed-optimization task, Claude went from about a 3× speedup in May 2025 to about 52× in April 2026; a skilled human reaches about 4× in four to eight hours. In an open-ended AI safety research problem, Claude agents closed 97% of a performance gap that two human researchers closed 23% of in a week, using about $18,000 of compute, though humans chose the problem and the scoring.
OpenAI said in February 2026 that early versions of GPT-5.3-Codex helped debug its own training and evaluation. Narrower self-improving systems, such as DeepMind's AlphaEvolve and the Darwin Gödel Machine, improve code or prompts against a fixed score they do not control. Both Anthropic and OpenAI say fully autonomous RSI is not happening yet. The remaining human role is choosing which problems matter and judging which results to trust.
Expected trajectory
METR measures the length of tasks AI can complete on its own; that length has been doubling about every four months. Anthropic projects that tasks taking a skilled person days could come into range in 2026, and tasks taking weeks in 2027. It describes three possible futures:
| Scenario | What happens | Anthropic's view |
|---|---|---|
| Stall and diffuse | Progress flattens, but today's models still spread through the economy | Possible, not likely |
| Compounding efficiency | AI does most of the work, humans set direction and verify. A 100-person firm does the work of thousands | The likely path |
| Full RSI | AI designs and trains its successors; humans move to oversight | Plausible, timing and safety uncertain |
Even in the faster scenarios, bottlenecks move rather than disappear. Anthropic already finds human code review is its new constraint. For a business, the equivalent is approval, verification and payer response times, not drafting speed.
What this means for your RCM operation
You can run a bounded, safe version of the same loop today. Agents propose changes to their own playbooks and prompts. Each change is scored against a fixed test set you control (the 300 graded cases from Module 12). Only changes that score higher, and that a human approves, are kept. Corrections feed the next LoRA or GRPO round. That is recursive improvement of your agent team, with the scoring and the final say kept in human hands.
Not to be confused with recursive language models, a separate long-context technique in which a model breaks a huge document into pieces and calls itself on each piece. Useful for 300-page payer manuals, but unrelated to self-improvement.
Quiz
Tap an answer to see the explanation. Your progress is saved in this browser.