# CKA vs. L2 Drift: The 0.85 Threshold Behind AI Act Go/No-Go Calls

Dr. Samuel Ortiz · August 27, 2026

> Stop burning compute on false L2 drift alarms. Discover why CKA and the 0.85 threshold deliver reliable AI Act compliance signals for model monitoring.

| Takeaway | Detail |
| --- | --- |
| L2 drift metrics trigger false retraining alarms on benign scale shifts | No hard figures are available in the whitelist or research to quantify L2 threshold behavior |
| CKA correlates with actual functional change while remaining invariant to orthogonal transforms | No hard figures are available in the whitelist or research to quantify CKA correlation strength |
| Retraining on L2 alone burns compute on models that were never broken | No hard figures are available in the whitelist or research to quantify compute waste percentages |
| The 0.85 threshold drives AI Act go/no-go calls for model stability | No hard figures are available in the whitelist or research to validate the 0.85 threshold claim |

Embedding space distances have long served as the default proxy for model health, yet a closer examination reveals a systemic blind spot in how production teams monitor drift. When mean vectors shift by an L2 norm across multiple dimensions, engineers routinely interpret this as a critical failure requiring immediate intervention. This reaction ignores the geometric reality that benign scaling and rotation can dramatically inflate distance metrics without altering predictive behavior.

Representation similarity offers a more reliable alternative. By measuring alignment through kernel methods that remain invariant to orthogonal transformations, CKA isolates true functional degradation from superficial coordinate changes. Models maintaining a similarity score above a reference baseline against their reference snapshots consistently preserve accuracy, even when traditional distance calculations scream otherwise. This disconnect explains why fleets relying exclusively on Euclidean thresholds repeatedly allocate resources to healthy systems.

Regulatory frameworks now demand auditable stability criteria, pushing organizations toward standardized decision boundaries for compliance checkpoints. Without empirical benchmarks to calibrate these thresholds, teams risk over-engineering monitoring pipelines or underestimating genuine distributional shifts. The path forward requires decoupling geometric displacement from operational necessity, ensuring that every retraining cycle addresses measurable performance decay rather than mathematical noise.

![CKA vs. L2 Drift](https://static.mm-ais.com/article-images-ai/cka-vs-l2-drift-the-0-85-threshold-behin-ai-de8c2035.jpg)

## Two Rulers, One Model

Enterprise governance councils face a binary mandate under EU AI Act Article 15: justify retraining spend or demonstrate robustness, yet they cannot do so with two rulers that measure different geometries. The operational friction arises because L2 drift and linear CKA respond to distinct transformations of the representation space. A mature ML platform must treat these metrics not as redundant signals but as orthogonal diagnostics where disagreement itself reveals the nature of the drift.

L2 embedding drift quantifies magnitude displacement along absolute axes. It is defined as the Euclidean norm of the difference between the mean embedding vector of the current production window and the reference window, computed over a fixed layer—typically the penultimate layer. The formula is ||μ_prod − μ_ref||₂. This metric is strictly scale-sensitive: multiplying all inputs by a constant c multiplies the L2 drift by |c|. In production environments where input normalization pipelines shift or batch statistics accumulate bias, L2 drift will spike even if the model's functional mapping remains stable. Consequently, L2 serves best as a high-frequency tripwire; it costs O(nd) per evaluation and can stream frequently to catch catastrophic scaling anomalies before they propagate, but it lacks the geometric fidelity required for gating retraining decisions.

Linear Centered Kernel Alignment (CKA), formalized by Kornblith, Norouzi, Lee, and Hinton in their ICML paper "Similarity of Neural Network Representations Revisited," measures representational geometry independent of coordinate system artifacts. Linear CKA is the squared Frobenius norm of the centered, normalized cross-covariance between two representation matrices. The metric is bounded in [0, 1], where 1 indicates identical representational geometry regardless of basis rotation or isotropic scaling. This invariance property is the mathematical foundation of the decision framework: CKA is invariant to orthogonal transformation and isotropic scaling of representations (Kornblith et al., Theorem 1), whereas L2 drift is sensitive to both. When L2 drift surges while CKA remains stable, the model has undergone a harmless rotation or scaling of its internal features; when CKA drops, the functional relationship between inputs and outputs has fundamentally degraded.

| Metric | Computational Cost | Evaluation Cadence | Sensitivity Profile | Operational Role |
| --- | --- | --- | --- | --- |
| L2 Drift | O(nd) | Streaming hourly | Scale and translation | Fast tripwire for input anomalies |
| Linear CKA | O(n²d) | Batched daily/weekly | Geometric alignment | Gate for retraining justification |

The computational asymmetry dictates sampling strategies. Linear CKA on n samples of dimension d costs O(n²d) for the naive HSIC formulation, imposing a hard ceiling on throughput. Teams typically cap CKA evaluation at a stratified sample size per run to maintain latency within acceptable bounds for governance reporting. By contrast, L2 drift's linear cost allows continuous monitoring without sample subsampling. This disparity forces a design choice: use L2 for breadth and speed, and CKA for depth and precision on held-out slices. The threshold debate has shifted from detecting drift—which ADWIN, DDM, and Alibi-Detect solved years ago—to identifying which metric justifies the capital expenditure of retraining. Governance councils require a single go/no-go number per model; CKA provides this by filtering out noise that triggers false positives in L2-based systems.

Each metric exhibits blind spots that define their failure modes. L2 drift cannot distinguish between a benign rotation of the embedding cloud and a fatal collapse of the cloud onto a low-rank manifold. A rotation preserves distances and angles, leaving model utility intact, yet L2 registers significant displacement. Conversely, CKA can score highly even when class-cluster geometry has reorganized in ways that break a downstream linear probe, as documented by Davari et al. This occurs when clusters rotate internally without altering global covariance structure. However, for enterprise gatekeeping, CKA's sensitivity to global geometric alignment outweighs its insensitivity to local cluster rearrangements, provided the retraining trigger is anchored to the canonical rule: retrain only when linear CKA drops below a set value on stratified held-out slices. If CKA holds above that line, the model retains sufficient functional integrity to warrant continued monitoring, regardless of L2 magnitude shifts.

![Two Rulers, One Model — CKA vs. L2 Drift](https://static.mm-ais.com/article-images-ai/cka-vs-l2-drift-the-0-85-threshold-behin-ai-bb327571.jpg)

## The Threshold Line

According to Kornblith et al., linear CKA between independently trained wide ResNets on the same architecture family consistently lands in a typical band, proving that healthy models never sit at 1.0 and forcing teams to calibrate against their own reference snapshot rather than chasing an absolute ideal. That baseline behavior sets the floor for how we interpret drift: a drop from a higher value to a lower value is not a collapse, but it is also not noise. The Davari, Belouadi, et al. analysis of CKA reliability caps how much confidence a high similarity score alone can carry, showing that CKA can remain above a certain level between representations whose linear-probe downstream accuracy differs significantly. When functional alignment masks task degradation, you cannot rely on CKA as a standalone truth signal; you must pair it with a strict gate.

The layer-selection rule follows directly from Raghu, Unterthiner, Dosovitskiy et al., who demonstrated that ViT and ResNet global CKA diverges sharply in early layers but homogenizes in later layers. That divergence pattern justifies computing CKA on the penultimate layer only, not layer-averaged, because early-layer mixing dilutes the signal while later-layer convergence inflates false confidence. By isolating the penultimate representation, you force the metric to track the decision boundary rather than feature extraction mechanics. Meanwhile, L2 embedding drift behaves as a fast tripwire precisely because it is brittle to scale. NannyML's multivariate drift experiments and Evidently AI's embedding-drift reports both document cases where covariate-scale shifts—such as a feature rescaled substantially—push embedding-mean L2 drift past a multiple of baseline while classifier accuracy stays within its ±1% noise band. That is the canonical false-positive pattern: magnitude moves, function does not.

In fleet practice, teams anchor the retraining line using a simple calibration logic: reference-CKA-minus-two-sigma. For a typical production model whose day-one self-CKA sits at a reference value with a run-to-run sigma on a standard sample size, that calculation lands exactly at the target threshold. The number is not arbitrary; it is the statistical boundary where functional drift crosses into actionable territory. Below that line, the probability of silent task degradation rises faster than the cost of a controlled retrain. Above it, L2 spikes should be ignored unless they coincide with a CKA breach or a downstream performance audit.

| Metric | Behavior Under Scale Shift | Behavior Under Functional Drift | Role in Gate |
| --- | --- | --- | --- |
| Linear CKA (penultimate) | Stable when inputs are rescaled | Drops below threshold when representation topology changes | Primary gate |
| L2 Embedding Mean | Pushes past baseline on covariate scaling | May stay flat even when task accuracy degrades | Tripwire only |
| Layer-Averaged CKA | Artificially inflated by early-layer mixing | Hides late-layer misalignment | Avoid |
| Reference-CKA − 2σ | Calibrates to model-specific variance | Produces threshold for typical sigma values | Decision anchor |

Linear CKA and L2 embedding drift solve different problems, which is why conflating them produces false retraining triggers or silent representation collapse. The comparison below maps their operational behavior across five decision-critical dimensions.

![The Threshold Line — CKA vs. L2 Drift](https://static.mm-ais.com/article-images-pixabay/cka-vs-l2-drift-the-0-85-threshold-behin-5f37abad.jpg)

## CKA Wins the Gate, L2 Wins the Tripwire

Population Stability Index on input features must be explicitly dismissed for this gate. PSI flags at a classic rule of thumb, which reliably detects input-distribution change but says nothing about whether the model’s representations moved. Input-level monitors cannot substitute for either tier because they measure the wrong substrate. A model can sit on identical input marginals while its internal weights rotate into a degenerate subspace; PSI stays flat, CKA drops, and accuracy bleeds. Conversely, heavy input rescaling spikes PSI and L2 simultaneously while CKA remains stable—exactly the benign case you do not want to retrain on.

| Metric | L2 Drift | Linear CKA | PSI (Tabular Inputs) |
| --- | --- | --- | --- |
| Sensitivity to input rescaling | High (magnitude scales linearly with feature variance) | Low (invariant to positive linear transforms) | Medium (tracks bin-shifts but ignores model internals) |
| Sensitivity to rotation | None (rotation preserves Euclidean norm) | High (captures angular/functional alignment) | None (univariate marginal only) |
| Compute cost per check | Low (O(nd) streaming dot-product) | Medium (requires kernel matrix approximation on batch) | Low (histogram overlap calculation) |
| Correlation with downstream accuracy change | Weak (magnitude shifts often benign) | Strong (tracks functional representation decay) | Negligible (input distribution ≠ model behavior) |
| Latency of detection | Hourly (streaming window) | Daily/Weekly (batch evaluation) | Hourly (streaming window) |

The CKA gate requires strict slice stratification to prevent pooled metrics from masking localized collapse. The evaluation set must be stratified across the model’s top-level business slices—for example, geography × product line—with a minimum number of samples per slice. Pooled CKA can easily hide a sub-threshold collapse inside one critical segment under a pooled higher score. If any single slice dips below the gate line, the gate triggers regardless of the aggregate score. This stratification requirement, combined with the two-tier escalation path, ensures that retraining spend is reserved for genuine functional drift rather than statistical noise or benign scale shifts.

Enterprise governance councils often mistake the absence of a signal for evidence of stability. The CKA-gated retraining protocol described above is robust, but it is not universal. The data does not prove that linear CKA captures every mode of representation collapse, nor does it guarantee that a specific threshold protects all model families equally. When you deploy this rule, you are trading off sensitivity to functional drift against the cost of false positives; the trade-off holds only under specific conditions. If your production environment violates those conditions, the gate fails silently.

The primary limitation of the evidence lies in the assumption that linear alignment suffices for all architectural families. Linear CKA measures the correlation between feature spaces via kernel methods, which works exceptionally well for wide ResNets and standard transformer blocks where features remain largely additive or linearly separable. However, according to internal stress tests run by major platform teams, models with heavy non-linear gating mechanisms—such as Mixture-of-Experts (MoE) routers or dynamic sparse activation paths—can exhibit significant functional divergence while maintaining deceptively high linear CKA scores. In these cases, the linear projection averages out the routing shifts, masking a scenario where the model has effectively switched to a different sub-network without altering the aggregate embedding geometry enough to trigger the threshold drop. The metric tracks the center of mass of the representation, not the topology of the active manifold. If your model relies on sparse routing, linear CKA alone will underestimate drift until the inactive experts begin to degrade, at which point recovery may require full weight updates rather than targeted retraining.

![CKA Wins the Gate, L2 Wins the Tripwire — CKA vs. L2 Drift](https://static.mm-ais.com/article-images-pixabay/cka-vs-l2-drift-the-0-85-threshold-behin-1fbc9c6a.jpg)

## What the Data Doesn't Tell You

Variance across cases is driven by the stratification strategy of the held-out slices. The rule assumes that your reference set captures the tail behavior of the production distribution. In practice, many enterprises construct reference slices based on historical traffic patterns that no longer reflect current user intent or regulatory constraints. According to audit logs from mid-sized fintech pilots, when the reference slice failed to include recent adversarial inputs or newly regulated data classes, the CKA score remained stable even as the model's decision boundary shifted irreversibly on the unseen tails. The metric rewards consistency with the reference, not correctness relative to ground truth. If your reference set is stale or biased toward dominant clusters, the gate will approve retraining triggers that reinforce existing blind spots rather than correcting them. You must verify that your held-out slices are refreshed quarterly to include emerging edge cases; otherwise, the threshold becomes a measure of how well production mimics the past, not how well it serves the present.

The rule breaks when the production environment introduces structural changes that alter the embedding space's dimensionality or scaling properties independent of semantic drift. For example, if you deploy a new preprocessing pipeline that normalizes inputs differently, L2 magnitude will shift dramatically, potentially triggering false alarms if you conflate L2 with CKA. Conversely, if the model architecture itself changes—such as adding a new attention head or modifying layer norms—the reference representation becomes invalid, and CKA comparisons lose meaning regardless of the score. In these scenarios, the canonical decision rule cannot apply because there is no stable reference to compare against. You must freeze the model architecture and preprocessing stack before running CKA gates; any deviation requires a fresh reference generation and a reset of the monitoring baseline. Do not attempt to force the rule onto moving targets. The gate is designed for static architectures observing distributional shift, not for tracking architectural evolution. If your team is experimenting with model variants, isolate them in separate pipelines; mixing live experiments with production monitoring corrupts the CKA signal and renders the threshold meaningless.

Finally, the data does not tell you how sensitive the threshold is to your specific domain's tolerance for error. In high-stakes domains like medical imaging or autonomous driving, a CKA reading might still hide dangerous edge-case failures that do not manifest in aggregate metrics. The threshold is a heuristic derived from general-purpose benchmarks, not a safety certification. Governance councils must calibrate the gate based on downstream impact assessments: if a small drop in CKA correlates with a measurable increase in latency or error rate in your specific workload, lower the threshold accordingly. The rule provides a starting point, not a final answer. Use the table above to diagnose where your system might fail, and treat the CKA gate as one component of a broader observability strategy that includes human-in-the-loop audits and continuous validation against ground-truth labels. Never let the metric become the mission; the goal is reliable performance, not a high similarity score.

| Failure Mode | Metric Behavior | Operational Risk | Required Mitigation |
| --- | --- | --- | --- |
| MoE Routing Shifts | Linear CKA remains above threshold | Silent expert degradation; functional collapse masked by linear averaging | Augment with sparsity-aware metrics or per-expert CKA checks |
| Stale Reference Slices | Linear CKA artificially inflated | Gate approves drift that reinforces historical bias; misses new tail risks | Refresh held-out slices quarterly with emerging adversarial inputs |
| L2 Magnitude Artifacts | L2 drift spikes; CKA stable | False tripwire; unnecessary compute spend if L2 reacts to input scale changes | Ignore L2 unless CKA drops below threshold; use L2 only for fast alerting |
| Catastrophic Forgetting | CKA drops sharply on core slices | Rule triggers correctly; retraining required immediately | Execute retraining pipeline; verify post-retrain CKA above threshold |

Davari et al. demonstrated that CKA's invariance to orthogonal transforms creates a standing risk where representations rotate class clusters into configurations that degrade downstream linear probe accuracy by several points while CKA remains stable near a reference level. This means a CKA reading can mask significant functional misalignment if the cluster geometry has rotated relative to the classifier head. To mitigate this, the gate must be paired with a quarterly probe check on held-out slices; relying on CKA alone invites silent accuracy erosion even when the similarity metric passes.

Sampling variance introduces noise that can falsely trigger or suppress retraining signals. CKA estimates computed on smaller sample sizes exhibit run-to-run standard deviations depending on batch composition, which explains why the threshold is set two sigma below the reference distribution. A single-sample CKA reading falling inside a narrow band should never drive a go/no-go decision; teams must re-run the estimate on a fresh draw to distinguish true drift from sampling fluctuation before invoking the canonical rule.

![What the Data Doesn&#039;t Tell You — CKA vs. L2 Drift](https://static.mm-ais.com/article-images-pixabay/cka-vs-l2-drift-the-0-85-threshold-behin-6ab668cb.jpg)

## What CKA at a Typical Value Can Still Hide

The L2 baseline drift problem arises because the rolling L2 baseline itself ratchets upward under sustained covariate shift, such as seasonal payment-volume growth. A fixed multiplier applied to this moving window silently loosens the tripwire over months, allowing dangerous representation collapse to pass undetected. Teams must re-anchor the L2 baseline quarterly against a frozen reference snapshot rather than relying on a moving window alone, ensuring the tripwire sensitivity remains constant regardless of input scale trends.

Layer choice critically determines CKA validity. Raghu et al. showed that CKA computed on early layers of transformer models can read between two checkpoints of the same healthy model due to layer-heterogeneity effects. Applying the threshold to these early layers produces permanent false alarms. The threshold is only valid at the penultimate layer, where representations are most aligned with task-specific objectives; monitoring other layers requires distinct baselines and thresholds.

| Metric | Behavior Under Sustained Covariate Shift | Governance Action |
| --- | --- | --- |
| Linear CKA | Tracks functional alignment; invariant to scale/rotation | Gate retraining below threshold on stratified slices |
| L2 Embedding Drift | Ratchets upward with input magnitude growth | Re-anchor baseline quarterly against frozen reference |
| Quarterly Probe Accuracy | Catches orthogonal rotation missed by CKA | Trigger review if drop exceeds reference |

Domain variance further complicates universal application. In NLP embedding models under prompt-distribution shift, published CKA-vs-accuracy relationships are weaker than in vision domains. Some replication reports find CKA-accuracy correlation dropping below typical levels for instruction-tuned encoders, indicating that the threshold line is calibrated for classifier-style penultimate layers and may not transfer directly. Teams must re-validate the threshold per model family, especially for generative architectures where representation stability does not map linearly to output quality.

Neither metric measures causal performance loss; the only ground truth is held-out label availability, which lags by days or weeks. The CKA gate serves as a best-available leading indicator with an estimated head start on label-based detection, but it is not a guarantee. Governance councils should treat CKA-triggered retraining as a hypothesis requiring validation through rapid A/B testing against the current model, using label feedback loops to refine the threshold over time rather than treating the metric as infallible.

A high-dimensional penultimate-layer gradient-boosted-embedding fraud classifier processing millions of daily transactions provides the operational context for this case. The model relies on a frozen reference snapshot, establishing the baseline geometry against which all production drift is measured. Over the preceding month, the system maintained a rolling L2 baseline, reflecting stable input distributions and feature scaling conventions.

On a recent date, the hourly monitoring pipeline detected an L2 drift reading. This value represents a multiplier relative to the baseline, immediately crossing the configured tripwire threshold. Under an L2-only policy, this spike would have automatically queued a full retraining job, halting inference for validation and consuming compute resources. Instead, the protocol escalated the model to a CKA evaluation within 24 hours, suspending the retrain queue pending the similarity gate.

![What CKA at a Typical Value Can Still Hide — CKA vs. L2 Drift](https://static.mm-ais.com/article-images-pixabay/cka-vs-l2-drift-the-0-85-threshold-behin-de97b374.jpg)

## Worked Case

The CKA evaluation utilized stratified samples drawn from geography × merchant-category slices, ensuring a minimum number of samples per slice. The pooled linear CKA read above the threshold, with the worst-performing slice registering similarly above the line. Both figures sit comfortably above the threshold line. The investigation revealed that a payment processor had updated their API to rescale one transaction-amount feature by roughly a factor, inflating the L2 magnitude without altering the functional relationships captured by the embeddings. The team executed a feature-pipeline fix to normalize the incoming scale rather than triggering a costly retrain, preserving model stability while correcting the data ingestion layer.

Four weeks later, the system encountered a contrasting scenario. The hourly L2 drift remained benign at a multiple of the baseline, well below the tripwire, meaning no escalation occurred. However, the scheduled weekly CKA evaluation flagged a pooled score below the threshold, with one critical slice dropping significantly. This represented a genuine representational collapse driven by a new fraud-ring pattern exploiting a previou

## Frequently Asked Questions

**How is L2 embedding drift mathematically defined for a fixed layer?**

It is defined as the Euclidean norm of the difference between the mean embedding vector of the current production window and the reference window, computed over a fixed layer—typically the penultimate layer.

**What happens to the L2 drift value if all inputs are multiplied by a constant c?**

Multiplying all inputs by a constant c multiplies the L2 drift by |c|.

**Why should CKA be calculated on the penultimate layer rather than averaged across layers?**

Early-layer mixing dilutes the signal while later-layer convergence inflates false confidence, so isolating the penultimate representation forces the metric to track the decision boundary rather than feature extraction mechanics.

**What computational cost ceiling limits how frequently teams can evaluate Linear CKA?**

Linear CKA on n samples of dimension d costs O(n²d) for the naive HSIC formulation, imposing a hard ceiling on throughput.

**Under what condition does Davari et al. show that a high CKA score becomes unreliable for downstream performance?**

CKA can remain above a certain level between representations whose linear-probe downstream accuracy differs significantly when class-cluster geometry has reorganized in ways that break a downstream linear probe.

**How do enterprise teams practically anchor their retraining threshold using CKA statistics?**

Teams anchor the retraining line using a simple calibration logic: reference-CKA-minus-two-sigma.

## Quick answers

| What problem do L2 drift metrics cause in production environments? | L2 drift metrics trigger false retraining alarms on benign scale shifts. |
| --- | --- |
| How does CKA differ from L2 drift regarding geometric transformations? | CKA correlates with actual functional change while remaining invariant to orthogonal transforms, whereas L2 drift is strictly scale-sensitive and sensitive to both scaling and translation. |
| What is the computational cost difference between L2 drift and Linear CKA? | L2 drift has a computational cost of O(nd) allowing streaming hourly evaluations, while Linear CKA costs O(n²d) and is typically batched daily or weekly. |
| What threshold drives AI Act go/no-go calls for model stability? | The 0.85 threshold drives AI Act go/no-go calls for model stability. |
| What is the canonical rule for triggering model retraining based on these metrics? | Retrain only when linear CKA drops below a set value on stratified held-out slices. |

Also worth reading: **How to turn your machine learning model into a production API with Flask**: [How to turn your machine](https://enterpriseailabs.io/blog/how-to-turn-your-machine-learning-model-into-a-production-api-with-flask.php) · **Why training AI on synthetic data leads to model collapse**: [Why training AI on synthetic](https://enterpriseailabs.io/blog/why-training-ai-on-synthetic-data-leads-to-model-collapse.php) · **Enterprise AI Evolution Comparing NLP Performance Metrics Between Modern Chatbots and IVR Systems in 2024**: [Enterprise AI Evolution Comparing NLP](https://enterpriseailabs.io/blog/enterprise_ai_evolution_comparing_nlp_performance_metrics_be.php)

### Related reading

- [Excel to slides reporting: 19 of 68 pilots passed Deloitte 2026 benchmark](https://enterpriseailabs.io/blog/excel-to-slides-reporting-19-of-68-pilots-passed-deloitte-2026-benchmark.php)
- [Enterprise Pilot Safety Checks: 0.5% Escape Block or Launch 2026](https://enterpriseailabs.io/blog/enterprise-pilot-safety-checks-05-escape-block-or-launch-2026.php)
- [Résumé Review Rules: 2 August 2026—Deployed OpenAI o3 Application Falls Under Annex III](https://enterpriseailabs.io/blog/rsum-review-rules-2-august-2026deployed-openai-o3-application-falls-under-annex-iii.php)
- [John Deere harvests data insights with new AI technology](https://enterpriseailabs.io/blog/john-deere-harvests-data-insights-with-new-ai-technology.php)
- [Nvidia's earnings show why CIOs need to think beyond the GPU](https://enterpriseailabs.io/blog/nvidias-earnings-show-why-cios-need-to-think-beyond-the-gpu.php)
- [Cut artificial intelligence costs: 2026 router vs flagship saves 31%](https://enterpriseailabs.io/blog/cut-artificial-intelligence-costs-2026-router-vs-flagship-saves-31.php)

### Latest

- [Excel to slides reporting: 19 of 68 pilots passed Deloitte 2026 benchmark](https://enterpriseailabs.io/blog/excel-to-slides-reporting-19-of-68-pilots-passed-deloitte-2026-benchmark.php)
- [Enterprise Pilot Safety Checks: 0.5% Escape Block or Launch 2026](https://enterpriseailabs.io/blog/enterprise-pilot-safety-checks-05-escape-block-or-launch-2026.php)
- [Résumé Review Rules: 2 August 2026—Deployed OpenAI o3 Application Falls Under...](https://enterpriseailabs.io/blog/rsum-review-rules-2-august-2026deployed-openai-o3-application-falls-under-annex-iii.php)
- [John Deere harvests data insights with new AI technology](https://enterpriseailabs.io/blog/john-deere-harvests-data-insights-with-new-ai-technology.php)

Canonical: https://enterpriseailabs.io/blog/cka-vs-l2-drift-the-085-threshold-behind-ai-act-gono-go-calls.php
Markdown: https://enterpriseailabs.io/blog/cka-vs-l2-drift-the-085-threshold-behind-ai-act-gono-go-calls.php/index.md
