| Takeaway | Detail |
|---|---|
| Compare both candidates at 20 concurrent requests. | Use the same model, prompt mix, context limits, and 20-request load on GPU and CPU. |
| Qualify the GPU only when the full workload fits in VRAM. | Any VRAM overflow or CPU spill disqualifies the GPU configuration. |
| Require the GPU to pass both p95 latency and error-rate gates. | The controlled 20-request test must meet the defined p95 latency SLO and error-rate limit. |
| Choose CPU if the GPU fails any qualification gate. | CPU is the defensible choice when latency, error rate, VRAM capacity, or CPU-spill requirements are not satisfied. |
This guide provides a controlled GPU-versus-CPU decision framework for 20 concurrent local-model requests. It identifies the memory, latency, and error-rate conditions that determine which deployment is more defensible.

Route requests before measuring hardware
Place a routing layer in front of separate model workers rather than choosing a device inside each worker. For every request, record its request ID, selected model, queue wait, prompt-token count, generated-token count, device, and completion status. Preserve those records as an event log so route decisions can be reconstructed after a latency spike, failure, or unexpected memory increase. Olla’s benchmarking guide recommends isolating setup from timed work, separating sub-benchmarks, and enabling memory and allocation profiling; the same discipline applies here, except the unit under test is an inference request rather than endpoint selection.
At concurrency 20, treat admission control as part of routing. Before accepting work for a target worker, verify that it can reserve memory for the model weights, the request’s KV cache, and runtime workspace. If the reservation would exceed the worker’s limit, the router must reject the request, defer it, or send it to an explicitly permitted fallback rather than starting execution and hoping the workload fits. A useful check is to compare the highest observed prompt and generated-token combination with the worker’s declared context and memory limits.
This section defines model switching as admission control plus queueing, not merely selecting a different model name. A switch occurs only when the router evaluates a new request’s resource requirements, accounts for work already admitted, and either places that request in the appropriate queue or changes its destination. Queue delay caused by a correct deferral should be visible in the request record, while worker saturation should be visible as a routing or admission outcome. That distinction exposes whether a “successful” model change actually improved service or simply moved contention elsewhere.
Exercise three switching paths independently. For warm switching between already-loaded models, keep both workers resident and measure queue wait and completion status without counting weight loading. For cold switching, include weight-reload time in the request’s queue wait and identify the reload in the event log. For overflow routing, deliberately make the preferred worker unable to admit a request, then verify that the router does not dispatch it there. Run each path with the same model, prompt mix, context limits, and request load so that the results reveal routing behavior rather than differences in workload.
Review each path for the same operational evidence: every request has one selected model and one recorded outcome, reserved capacity matches estimated demand, and rejected or deferred requests remain distinguishable from failed generations. Also confirm that warm, cold, and overflow results are not combined into a single average. If one path requires manual intervention, hides reload cost, or permits dispatch without a valid reservation, the router is not yet providing auditable model switching.

The available evidence sets a method
Olla’s benchmarking guide provides a reproducibility pattern rather than a local-model performance result. Its documented Go benchmark command and emphasis on separating setup, measuring concurrent performance, and profiling memory and allocations can guide the comparison. For this test, define the workload in advance, separate setup from timed execution, preserve the commands and environment, and record memory measurements for both candidates. Any endpoint, payload, or duration figures should be verified against the named guide before they are used in the local-model test.
Build the test around the actual service contract. Use the same model, prompt mix, context limits, request distribution, and 20 concurrent requests on both candidates. Record each request’s result and latency, then summarize p50 and p95 latency together with completion failures, timeouts, and errors. Keep queueing time distinct from model execution time so that a fast device cannot hide a slower admission path. Save the raw request and response logs alongside the test configuration; otherwise, a later comparison cannot establish whether both candidates were exercised under equivalent conditions.
For each candidate, document the host and execution environment rather than relying on a machine label or an unrelated benchmark artifact. Record CPU topology, accelerator model and memory capacity, relevant runtime or driver versions, model placement, scheduling behavior, and any memory-pressure, overflow, or spill events. Verify the exact environment fields in the test logs; none of those records should be treated as evidence of local-model inference speed by themselves.
Compare the candidates only after the runs are complete. The GPU candidate passes the hardware gate when the full workload fits in VRAM, the controlled 20-request test meets the p95 latency SLO and error-rate gate, and the logs show no VRAM overflow or CPU spill. If any of those conditions fails, the CPU deployment is the more defensible choice for predictable service, even if isolated GPU requests appear faster. The available sources do not establish a GPU-versus-CPU throughput number, so any claimed advantage must come from the recorded test results under this shared method.

GPU wins only after the gate passes
At a load of 20 concurrent local-model requests, hardware selection should be treated as a controlled acceptance test, not a preference based on device type. Run the same model, prompt mix, context limits, and request pattern on both candidates. The GPU is the explicit winner only if the entire workload stays on the GPU path, no memory-overflow events occur, and the measured p95 latency and error-rate gates pass without CPU spill. Otherwise, the CPU deployment is the more defensible choice for predictable service.
| Option | Strength | Failure mode at 20 concurrent requests | Decision |
|---|---|---|---|
| GPU | High parallelism and resident weights | VRAM exhaustion, KV-cache spill, or cold-load stalls | Explicit winner when memory and SLO gates pass |
| CPU | Larger system memory and simpler residency | Lower token throughput and queue growth | Explicit winner when GPU gates fail or CPU p95 is the only passing result |
For the GPU test, verify that all 20 requests are admitted to the intended execution path. Treat any fallback, offload, or request that requires CPU execution as a failed GPU gate, even if the aggregate result appears acceptable. Monitor VRAM use throughout the run, including model residency, KV-cache allocation, and any reload behavior. A run that finishes successfully but overflows memory or spills requests is not evidence that the GPU configuration is safe.
Evaluate the service objective directly: compare the GPU’s measured p95 latency with the CPU’s result, and compare both error rates with the defined acceptance gate. Also check whether queue time grows as requests wait for scarce device capacity. The Olla benchmarking guide, authored by Thushan, identifies memory profiling, concurrent performance, and performance goals as parts of a serious benchmark process; those checks help make this hardware decision auditable rather than anecdotal.
The decision rule is intentionally asymmetric. Treat the GPU as the selected deployment only when the full workload remains within VRAM, no request spills to CPU, and the controlled test passes the defined p95-latency and error-rate gates. If the GPU fails any of those conditions, the CPU is the more defensible fallback for predictable service, although the CPU must still pass its own service gates before it can be described as a passing deployment.

Count memory, queue time, and reload cost
To make capacity auditable, calculate resident memory as the sum of model weights, runtime workspace, and KV-cache memory for active requests. Measure the high-water mark during the full 20-request run, rather than relying on startup usage, because the live-cache term can rise as requests wait for service and generate tokens. Record VRAM headroom at that peak, and treat any overflow or forced movement to host memory as a failed configuration. The Olla benchmarking guide’s memory-profiling and allocation-analysis practices support separating setup from timed execution; its documented command, go test -bench=. -benchmem ./..., is an example of collecting benchmark and allocation data, not evidence of a particular inference result.
Measure end-to-end latency as queue wait plus prefill time plus decode time plus switching or reload time. Publish the components separately, together with the total, so a fast decode phase cannot conceal a long admission wait or a cold-model penalty. For every request, retain the request ID, selected model, device, prompt-token count, generated-token count, queue wait, prefill duration, decode duration, switching or reload duration, completion status, and total latency. This makes a 20-request test reviewable: a reviewer can determine whether the service met the p95 latency gate because inference improved, or merely because queueing and reload costs were moved elsewhere.
Use a cost worksheet that lists hardware purchase or rental cost, power cost, host and storage cost, software and observability cost, operational labor, and the cost of failed or retried requests. Enter measured quantities beside each item rather than filling in a device label. The worksheet should also show the number of workers, resident memory at the high-water mark, peak VRAM, average and p95 queue wait, average and p95 total latency, error rate, and any host-memory spill. Report the measurement window and whether the model was already resident; otherwise, reload frequency can be mistaken for steady-state capacity.
The central rule is simple: accept the GPU only when the measured workload fits in VRAM, remains within the memory limit through the concurrency peak, and passes the latency and error-rate gates without CPU spill. If the GPU incurs reloads, queue delays, or memory movement that pushes results outside those gates, the worksheet shows the cause and favors the CPU deployment for more predictable service.

Know where the decision rule breaks
The default decision rule breaks when an assumption used to justify the GPU stops holding under the measured workload. A device is not the right choice merely because it is a GPU; the selection remains defensible only while memory, queueing, switching, and request behavior continue to support the original conclusion. For this section, the rule breaks when the workload can be characterized by one of the conditions below and the GPU’s relevant failure mode appears in the controlled test.
| Condition | Rule breaks when | Rule still wins when |
|---|---|---|
| Multiple models | Switching is cold and reload time dominates token generation. | Every candidate model is preloaded or switching is rare. |
| Long contexts | KV-cache growth causes GPU overflow. | The measured peak remains below the VRAM reservation. |
| Bursty traffic | Queue wait dominates device compute. | The same burst profile is replayed on both devices. |
Check the model-switching path separately from steady-state generation. Record model-load time, queue wait, time to first token, and total completion time for each model transition. If reload time consumes most of the request’s service time, the GPU’s compute advantage is no longer the dominant fact. Preloading removes that particular cost, but it also makes the VRAM reservation larger, so the memory test must use the complete set of resident models rather than a single model in isolation.
For long prompts, measure memory at the longest supported context and the largest configured batch, then inspect the observed peak against the reservation. KV-cache growth is the relevant warning signal: a run that fits with short prompts can fail when concurrent requests retain longer histories. Check not only whether the process starts successfully, but whether peak usage leaves enough headroom for the measured workload and whether any request is forced into CPU spill.
Treat batch size, prompt length, output limit, quantization, sampling parameters, and the scheduling policy as workload controls, not incidental settings. Change one factor at a time, preserve the prompt mix, and replay the same burst pattern on both candidates. The Olla benchmarking guide describes sub-benchmarks, payload-size tests, memory profiling, and allocation analysis; those practices are useful for making each condition observable, but they do not establish a universal GPU-versus-CPU result. The practical edge-case check is therefore simple: identify the condition, reproduce it, and retain the GPU choice only if the full comparison still satisfies the service gates.

Use a reproducible 20-request worksheet
Test one local model with exactly 20 concurrent requests on each candidate. Replay the identical requests, model version, prompt mix, context limits, output limits, and sampling settings so that GPU and CPU results remain comparable. Assign each request a unique ID, and run both warm and cold-switch scenarios separately so model-loading time is not hidden in steady-state execution. Treat the context and output limits as controlled test settings rather than universal model limits.
Before execution, preserve the model version, quantization or checkpoint, runtime version, device configuration, scheduler settings, context limits, and defined latency and error-rate gates. Keep the prompt set, request order, concurrency of 20, output limit, and measurement procedure unchanged between GPU and CPU runs. Olla’s benchmarking guide can inform repeatable test structure, but its Go benchmark commands are not evidence of performance for this local-model workload.
| Checkpoint | GPU result | CPU result | Pass condition |
|---|---|---|---|
| C1: resident-memory high-water mark | Record GB | Record GB | Below the device reservation, with no VRAM overflow or CPU spill during the run |
| C2: p95 end-to-end latency at concurrency 20 | Record milliseconds | Record milliseconds | At or below the defined SLO |
For each run, retain request-level measurements and calculate p95 using the same end-to-end latency definition on both devices. Start the clock at request admission and stop it when the final generated token becomes available, thereby including queueing, prefill, decode, and any applicable switching or reload time. Preserve cancellations, timeouts, retries, failures, and reload events instead of replacing them with successful completions. A run with missing measurements remains unevaluated, while a run that violates a defined gate is marked as failed and its logs are retained.
Compare the two columns only after both warm and cold-switch runs use the same fixture. A blank cell is an unmeasured result, not a pass. If the GPU misses either the memory or latency condition, or if it shows VRAM overflow or CPU spill, do not treat the run as a qualifying GPU result. If both candidates pass, retain the measured results as a controlled comparison and make the selection from the recorded gates. This worksheet supplies the test record; it does not turn a device label into evidence of service performance.
Apply five final switching rules
Rule 1: Approve GPU routing only after the full gate passes. Run the identical model, prompt mix, context limits, and 20-request load on the GPU and CPU. Approve GPU routing only if all 20 requests remain resident without CPU spill, VRAM overflow events equal zero, p95 latency meets the service SLO, and the error rate stays within its configured gate. If any condition fails, do not label the GPU the winner. This section converts the controlled test into five deployment actions: approve routing, manage cold starts, compare cost, protect repeatability, and record the decision.
Rule 2: Treat cold switching as a separate deployment test. Measure the time from a switch request to the point where the GPU can accept the full workload safely. If that cold-switch time exceeds the service’s allowable reload budget, preload the required models before accepting traffic or route that workload to CPU. Do not use warm-run latency to approve a cold-switch procedure, because warm-run results omit the loading and initialization delay that production switching may incur.
Rule 3: Use cost as the tiebreaker only after both candidates pass. If the GPU and CPU both satisfy the p95 latency and error-rate gates, calculate total test cost divided by the number of successful requests, then compare the resulting cost per successful request. The calculation must use the same workload and counting window on both candidates. A lower GPU cost does not compensate for an SLO failure, and a higher GPU cost does not justify retaining it once its reliability and latency gates have passed.
Rule 4: Repeat the controlled test after deployment changes. Re-run the comparison whenever the model, prompt mix, context limit, concurrency load, runtime configuration, or hardware allocation changes. Keep the request accounting consistent so queue time, generated tokens, failures, and retries do not disappear between runs. Olla’s benchmarking guide supports separating setup from timed work with b.ResetTimer(); the same discipline applies here: prepare the model, exclude setup where appropriate, and measure the operational path rather than installation state.
Rule 5: Record one auditable routing decision and a fallback. Store the chosen device, test configuration, p95 latency, error rate, overflow count, CPU-spill count, cold-switch time, and cost per successful request. Mark GPU deployment as approved only when every required gate passes; otherwise select CPU routing for predictable service. If later evidence shows a failed gate, rerun the same test after remediation and retain the previous result as the deployment audit trail.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Run the same local AI model on GPU and CPU using the same prompt mix, context limits, and 20-request load. | A controlled workload makes the deployment comparison defensible. |
| 2 | Confirm that the full GPU workload fits in VRAM before evaluating performance. | The GPU qualifies only when the entire workload remains in VRAM. |
| 3 | Reject the GPU configuration if any request overflows VRAM or spills to CPU. | Either condition disqualifies the GPU deployment. |
| 4 | Measure the controlled test against the defined p95 latency SLO and error-rate limit. | The GPU must pass both service gates; strong throughput alone is insufficient. |
| 5 | Choose the GPU only if it passes the p95 latency and error-rate gates with no VRAM overflow. | Meeting every qualification gate makes GPU deployment the valid choice. |
| 6 | Choose the CPU if the GPU fails any latency, error-rate, VRAM-capacity, or CPU-spill requirement. | The CPU is the defensible fallback when the GPU cannot satisfy the complete workload. |
Frequently Asked Questions
What workload must be held constant when comparing GPU and CPU candidates?
Use the same model, prompt mix, context limits, and 20-request load on both GPU and CPU.
Can a GPU remain qualified if part of the workload spills into CPU?
No; any VRAM overflow or CPU spill disqualifies the GPU configuration.
Which performance gates must the GPU pass under the controlled load?
The GPU must pass both the p95 latency SLO and the error-rate limit.
When should CPU be selected instead of GPU?
Choose CPU if the GPU fails any qualification gate.
Where should request routing occur for separate model workers?
Place a routing layer in front of separate model workers rather than choosing a device inside each worker.
Which practices should be applied when benchmarking the deployment?
Isolate setup from timed work, separate sub-benchmarks, and enable memory and allocation profiling.
Quick answers
| How should GPU and CPU candidates be compared for 20 concurrent requests? | Compare both candidates at 20 concurrent requests using the same model, prompt mix, context limits, and 20-request load on GPU and CPU. |
| When is a GPU configuration disqualified? | Any VRAM overflow or CPU spill disqualifies the GPU configuration. |
| What should happen if the GPU fails a qualification gate? | Choose CPU if the GPU fails any qualification gate. |
| Where should routing occur for a local AI deployment? | Place a routing layer in front of separate model workers rather than choosing a device inside each worker. |
| What should be recorded for every request? | For every request, record its request ID, selected model, queue wait, prompt-token count, generated-token count, device, and completion status. |
Also worth reading: Driving superior enterprise AI performance with optimization algorithms: Driving superior enterprise AI performance · How to turn your machine learning model into a production API with Flask: How to turn your machine · Why training AI on synthetic data leads to model collapse: Why training AI on synthetic