跳到正文
GitHub Blog · AI & ML· Erin Havens·· 4 小时前AI 评分54

GitHub 推出基于 ModernBERT 的 AI 通用密钥检测模型扩展 push protection

Secret protection must scale with software

AI 导读

GitHub 发布用于 push protection 的微调 ModernBERT 分类器,可在上下文中评估候选密钥、单批耗时不足 2 毫秒,使可阻止的密钥数量有望翻倍以上,功能目前处于 private preview。

正文 · 原文

Developers aren’t becoming more careless; they’re being outpaced. The tools that let developers create more software should also take on more of the work of protecting it.

October 7, 2026

|

7 minutes

  • Share:

Today, one in three pull requests on GitHub involves an AI agent. A year ago, that number was fewer than one in 10. If that pace holds, within the next two years, most of the code pushed to GitHub could be written by an agent. Much of it may never be fully read by a human.

If developers and agents move faster, we have a responsibility to ensure protection keeps up with the accelerated rate of code creation. That means preventing more leaks before they happen and making the response to exposures that remain less dependent on manual human effort.

This is a pivotal point for leaked secrets. Developers aren’t becoming more careless; they’re being outpaced. The tools that let developers create more software should also take on more of the work of protecting it.

In this essay, I share the nine quarters of data behind that claim. I also introduce the fine-tuned classifier we built with Microsoft Applied Sciences to extend push protection to unstructured secrets. The model assesses a whole set of candidate secrets in less than two milliseconds and could more than double the number of secrets that we can prevent.

Outpaced, not careless

A new secret appears in publicly visible code about once every two seconds, doubling yearly for the past three years. Public discourse is quick to jump to the idea that AI made developers careless.

Between Q2 2024 and Q2 2026, screened pushes grew 2.84 times while pushes carrying credentials grew 2.59 times. Across nine complete quarters of data, we found no statistically detectable trend regarding per-push prevalence. At the same time, we found data suggesting that, more than ever, developers understand the risk of accidental exposures and are less willing to accept that risk. Over the same period, the share of push-path blocks overridden by developers fell linearly from 6.63% to 3.93%. These figures challenge the common claim that agents are causing developers to become more careless.

2026 Q2 · 574M pushes · with secrets

Public pushes, Q2 2024–Q2 2026. Push prevalence is the share with a detected secret. Covers supported provider patterns, including GitHub’s own tokens.

At a fixed rate, doubling activity doubles expected exposures. If each exposure requires the same human response, the workload doubles too. The mean time to manually revoke a secret hovers around 40 days; roughly one in five took more than 90 days. We’re accelerating the creation of software while exposed credentials can remain usable for weeks or months, because human remediation can’t scale at the same pace as development.

Telling developers to be more careful cannot, on its own, solve that problem. As the amount of code grows, we must prevent more exposures and reduce human effort required by those that remain, if software development is to remain sustainable.

Prevention scales with compute

I’ve spent the past few years working on secret scanning at GitHub and the past year as product lead for the area. Our greatest impact has come from connecting the dots between detection to systems that can act.

GitHub’s catalog covers more than 150 technical partners through our secret scanning partnership program. Through our partner program, we work with participating secret issuers to build out detectors and report public exposures so they can respond. In Q2 2026, public scanning successfully reported an average of 26 credential matches a second, including repeat observations. Once notified, a large number of these partners immediately revoke the token: OpenAI API keys, Google Cloud account credentials, Slack webhooks, Hugging Face user tokens, SendGrid keys, etc. The owner may still need to replace the token, but revocation can happen without waiting for a developer to find and process a GitHub alert.

Push protection intervenes earlier. It stops recognizable credentials before it enters repository history, giving the developer or agent a chance to correct the change before there’s an exposure to investigate. We work with our technical partners to increase precision rates of their detectors as much as possible, until we’re confident enough to push-protect these secrets for the developer community by default.

Thanks to the efforts of our partners, in the past month, a secret was blocked by push protection at least once every second. When it comes to issuer-bound credentials, GitHub blocks more secrets than those which slip. I’m proud of how ordinary we’ve made that feel for developers.

When including additional secret types, push protection stops about 30% of newly detected secrets before they enter repository history. We find the remaining 70% after the credential is, unfortunately, already lost. And:

  1. Prevention scales with compute, but remediation still scales with people.
  2. Refusing a push costs compute; cleaning up a secret already lost to visible history costs a developer’s time and attention.
  3. As the amount of code grows, we must prevent more exposures and reduce the human effort required by those that remain, or else the volume of vulnerabilities introduced will become untenable.

Telling developers to be more careful cannot solve this imbalance. Recognizing more of these secrets, earlier in development flows, is work the platform must take on.

Solving the four-body problem

Before a secret crosses the push boundary, the cost of stopping it is small, and the decision is binary: block or allow. After it crosses, the same string can authenticate to a real system, and the cost is unbounded.

In many cases, our only detection clue may be the surrounding code and world context. A provider-issued token may have a recognizable prefix. An internal database password may be completely unstructured, with no identifying pattern at all. We were already using context to find these secrets post-push; the problem was balancing that context-aware judgement with other factors.

We refer to this as the “four-body problem” for secret protection: precision, latency, throughput, and cost are coupled constraints. Prevention must be worth a developer’s time. A finding suitable for later review may not justify blocking a push. A false positive interrupts a developer and makes the next block harder to trust. A check that is too slow, expensive, or difficult to scale limits how often it can run.

Protection at the push in under 2 ms

GitHub’s AI-powered generic secret detection model uses surrounding code context to block password-like values in a database URL, Kubernetes Secret manifest, and Dockerfile, while allowing the placeholder changeme.

Our new ModernBERT classifier assess candidate secrets in context, without generating code or prose. It’s not only more precise than existing LLM-based pipelines, but it’s incredibly fast, evaluating candidate batches in under two milliseconds. It’s also extremely cost efficient, enough to run at scale in the critical path.

The inclusion of our model in push protection makes it possible for us to more than double the number of secrets that we’re able to prevent. The feature is currently in private preview. Later this month, the feature will be available to organizations with GitHub Secret Protection across Enterprise Cloud and GitHub Teams. It will consume AI credits.

We’re also bringing the model to developer surfaces beyond the push.

  • Starting today, any organization with AI secret detection will be automatically updated to the new model. Alerts opened from these post-push scans remain included with an organization’s purchase of secret scanning at no additional cost.
  • The model will also ship with GitHub Enterprise Server 3.23 in public preview, bringing AI-detected alerts to Secret Protection customers even in air-gapped environments.
  • We’re adding the classifier to the /security-review command for the Copilot CLI and Copilot App, so Copilot users can address secrets before a push even without needing an organization’s GitHub Secret Protection plan. AI credit usage will be attributed to GitHub Secret Protection in your AI usage insights.

Looking forward

The future we want is one where developers can entrust more work to agents without supervising every request, and one where the number of people that an organization needs to keep its credentials safe no longer scales with the amount of code it writes. We owe the developer community the same progress in protecting software that we are delivering in producing it.

We want people to build more software. Our capacity to protect it should grow with our capacity to create it.

Written by

Erin Havens

Erin Havens is a Product Manager at GitHub, focused on security products. 100+ ships across products like Secret Protection and Dependabot (and counting).

Related posts

We do newsletters, too

Discover tips, technical guides, and best practices in our biweekly newsletter just for devs.

来源:GitHub Blog · AI & ML · github.blog