Google had a model rewrite a C library into Rust in a single prompt, then spent six days verifying it. What the harness caught, and why it decides production risk.
Google's security team published a pilot result that deserves closer attention than its headline received.
Per Google's own writeup on scaling memory safety, the team used Gemini to rewrite giflib, a roughly 3,000 line C image-processing library, into an ABI-compatible Rust replacement. The translation itself was a single prompt covering the complete codebase.
Then the actual engineering started.
According to InfoQ's analysis of the project, the process ran in three distinct stages, and only the first involved generation.
A single Gemini pass ported the full C logic to Rust. Google notes this approach was viable specifically because of the library's relatively small size, which is a constraint worth carrying forward rather than glossing over.
Human engineers reviewed the output and found unsound raw pointer semantics and incorrect Rust lifetime annotations.
This detail matters more than the headline. The model produced Rust, a language most engineers assume is memory-safe by construction, and parts of the output were not. AI-generated Rust containing unsound pointer handling is not safer than the C it replaces. It is a memory-safety bug wearing a safer language's syntax.
Critically, reviewers did not re-read every generated line evenly. They concentrated on FFI boundaries and lifetimes, which is where model output is most confidently wrong.
An automated differential fuzzer executed the original C and the new Rust side by side, feeding behavioural discrepancies back to the model for iterative patch synthesis.
The validation numbers are the point of the entire exercise:
| Validation stage | Scale | Duration |
|---|---|---|
| One-shot translation | ~3,000 lines of C | Single prompt |
| Human FFI and lifetime review | Targeted at unsafe seams | Not disclosed |
| Real-world regression corpus | 30M+ GIF files, bit-for-bit parity | Batch |
| Differential fuzzing | ~200M iterations, no functional drift | 6 days |
Generation was close to free. Confidence was expensive, and it was constructed rather than assumed.
Two outcomes justify the cost.
The differential fuzzer surfaced a pre-existing out-of-bounds write in Google's own legacy C patch. Not a defect the model introduced. A latent bug in production code, found because someone built apparatus capable of comparing two implementations under adversarial input.
Then CVE-2026-26740 was disclosed, an out-of-bounds heap write in the original C giflib requiring a fix across Linux distributions. Google's production nodes were already structurally immune, because that class of defect does not exist in the Rust implementation.
The usual objection to memory-safe rewrites, runtime overhead from mandatory bounds checks, did not materialise. Performance held, in part because removing the memory-safety risk allowed the team to eliminate the decoder sandbox previously wrapped around the C library.
Discussion on Hacker News and r/rust praised the fuzzing rigour while raising three objections that hold up.
One-shot translation likely does not scale. The approach worked partly because the library is small, a constraint Google states directly. Nobody has demonstrated single-pass LLM translation across a 300,000 line codebase, and several commenters argued deterministic transpilers such as c2rust remain more predictable at scale.
The effort ratio is undersold. Auditing subtle semantic regressions and fixing unsound FFI boundaries can exceed the generation step substantially, which means framing this as AI performing the rewrite understates the senior human time consumed.
Maintenance divergence is unsolved. Forking an upstream C dependency into a Rust reimplementation creates permanent reconciliation cost. Every upstream release becomes your integration problem.
All three are fair. None changes the transferable lesson.
Our read: this is the clearest public accounting anyone has published of what it actually costs to move AI-generated code onto a critical path.
Contrast it with how AI-generated code typically reaches production at a smaller company. Code is generated, it compiles, tests pass, someone skims the diff, it merges. The verification step is a human reading code under time pressure.
Same model quality in both cases. Entirely different risk profile. The difference is the harness.
This is the mechanism underneath the telemetry we covered in our analysis of the AI quality collapse, where incidents per merged pull request rose sharply as adoption scaled. The missing ingredient was never better models. It was verification capacity growing at the same rate as generation capacity.
A human reading a diff is attention, and attention does not scale with output. A differential comparison between old and new behaviour is mechanical, and it does.
Most teams will never fuzz a GIF decoder for six days. The transferable pattern is comparing two implementations mechanically across real inputs.
This pattern applies to any reimplementation: a service rewrite, a library swap, a database migration, or an AI-generated replacement of existing logic.
import hashlib
import json
from dataclasses import dataclass, field
@dataclass
class DiffReport:
checked: int = 0
mismatches: list = field(default_factory=list)
old_errors: int = 0
new_errors: int = 0
def normalise(value):
"""Canonicalise before comparing, or you will chase false positives
through dict ordering and float formatting for a week."""
return json.dumps(value, sort_keys=True, separators=(",", ":"), default=str)
def run_differential(corpus, old_impl, new_impl, max_report=50) -> DiffReport:
report = DiffReport()
for case in corpus:
report.checked += 1
try:
old_result, old_failed = old_impl(case), False
except Exception as exc:
old_result, old_failed = repr(exc), True
report.old_errors += 1
try:
new_result, new_failed = new_impl(case), False
except Exception as exc:
new_result, new_failed = repr(exc), True
report.new_errors += 1
# Error-path parity matters as much as happy-path parity.
if old_failed != new_failed:
if len(report.mismatches) < max_report:
report.mismatches.append({
"case": hashlib.sha256(repr(case).encode()).hexdigest()[:12],
"kind": "error_divergence",
"old": old_result,
"new": new_result,
})
continue
if normalise(old_result) != normalise(new_result):
if len(report.mismatches) < max_report:
report.mismatches.append({
"case": hashlib.sha256(repr(case).encode()).hexdigest()[:12],
"kind": "output_divergence",
"old": old_result,
"new": new_result,
})
return report
Two details carry the value. Normalising before comparison prevents a week of chasing dictionary ordering and float formatting as if they were real defects. And comparing error paths, not only successful returns, catches the divergence class that causes production incidents, since a reimplementation that raises where the original returned an empty result will pass every happy-path test you own.
A fixed corpus tests the inputs you thought of. Fuzzing tests the ones you did not.
import atheris # pip install atheris
import sys
def TestOneInput(data: bytes):
try:
old = old_impl(data)
old_failed = False
except Exception:
old, old_failed = None, True
try:
new = new_impl(data)
new_failed = False
except Exception:
new, new_failed = None, True
if old_failed != new_failed:
raise RuntimeError(f"error divergence on {data!r}")
if not old_failed and normalise(old) != normalise(new):
raise RuntimeError(f"output divergence on {data!r}")
atheris.Setup(sys.argv, TestOneInput)
atheris.Fuzz()
For Rust reimplementations specifically, cargo-fuzz with a libFuzzer target follows the same shape, calling into the original C through FFI and comparing both results inside a single fuzz target.
The harness is the deliverable. It is also the part small teams skip, and it decides whether AI-generated code becomes an asset or a liability.
It is burst-shaped work in the most literal sense. You need it intensely during a migration or rewrite, it requires someone who has built differential testing before and knows where reimplementations actually break, and it is not a standing role afterward. The harness continues working long after the engineer who built it has moved on.
This compounds the pipeline saturation problem described in our analysis of delivery bottlenecks and the surface area argument in our March edition. Generation capacity scaled. Verification capacity did not, and mechanical verification is the only kind that can catch up.
The engineers we embed at Percime Technologies build this specific apparatus: differential test infrastructure, fuzzing harnesses, regression corpora, and shadow deployment paths. The machinery that turns a plausible-looking rewrite into one you can defend in an incident review.
Google let a model rewrite a library in a single pass, then spent six days and 200 million iterations confirming the model had not quietly broken something.
The second part is the engineering.
© 2026 Percime Technologies. All rights reserved.