Skip to content
12 min read

Short Token Lifetimes Answer Only One of Three Ways Tokens Fail

NIST and CISA have finalized their token protection guidance. Its useful lesson is not one-hour lifetimes but a map of where tokens fail and which controls reach each failure.

Antonio J. del Águila

Knaisoma

It is tempting to reduce token security to a number: access tokens live for an hour, or fifteen minutes, or five. The number is worth having, but it answers a narrower question than token security asks. A short lifetime limits how long a stolen token stays useful. It does nothing when an attacker can mint tokens of their own, and nothing when a service accepts tokens it never properly checked.

On September 15, NIST and CISA released the final version of NIST IR 8587, Protecting Tokens and Assertions from Forgery, Theft, and Misuse. It is written for federal agencies and their cloud providers, but its authors say plainly that it applies to anyone using tokens in access management, and it is one of the most complete public statements so far of what good token handling looks like. Less than three weeks earlier, Kubernetes 1.37 made Pod Certificates generally available, giving workloads a built-in alternative to the bearer tokens they use today. Read together, they support a simple argument: tokens fail at three distinct points, each point needs its own controls, and many organizations have invested heavily in only one of them.

Three places a token can fail

A token passes through three hands. An issuer signs it, a holder carries it, and a verifier decides whether to believe it. Each hand is a separate failure point, and the attacker’s position is different at each.

Forgery happens at the issuer. If an attacker obtains a signing key, they no longer need to steal anyone’s token. They write their own, with whatever subject, scope and expiry they choose, and every verifier that trusts the key accepts it. The NIST announcement recalls the attack that shaped this work: foreign actors used forged tokens derived from a single stolen commercial signing key to reach agency email, taking more than 60,000 messages from one agency. Token lifetime was irrelevant to that attack, because the attacker controlled the lifetime.

Blind acceptance happens at the verifier. A service that does not verify signatures, algorithms and audiences will accept tokens nobody issued. On September 28, Help Net Security reported a researcher’s account of Titan, an internal Microsoft analytics service. According to that account, Titan checked tenant, audience and application claims one after another but never checked the signature, and accepted a token whose algorithm was set to none with an empty signature. Setting the user field to admin then matched a local account holding the Admin role. The researcher reported the flaw on September 5, and Microsoft locked the endpoint down on September 9 and paid a bounty. Microsoft had editorial control over the researcher’s write-up, so read the details as the researcher’s account rather than an audit. The lesson is the shape: issuer-side controls are wasted on a verifier that does not look.

Theft happens in transit and at rest with the holder. A legitimate token is copied from a log, a CI artifact, a compromised host or a peer service, and replayed from somewhere else. Short lifetimes were designed for this case and they help, but a stolen token is fully usable until it expires, and an attacker who keeps stealing can keep replaying.

What the final guidance asks for at each point

IR 8587 is organized around control enhancements rather than failure points, but its recommendations map cleanly onto the three.

1 hour

Recommended maximum validity for access and identity tokens

NIST IR 8587, sec. 5.3.1.1

90 days

Recommended maximum signing period on high-impact systems

NIST IR 8587, sec. 5.1.1.4

~250

Public comments incorporated into the final report

CISA, September 15, 2026

On the issuer side, the report asks for signing keys to be isolated according to risk, scoped to the lowest reasonable level such as a single tenant, and bound to one environment so that a development key cannot sign production tokens. Rotation should be automated, with overlapping acceptance of old and new keys, and the active signing period should be no more than 90 days on high-impact systems and under a year elsewhere. NIST notes that the key protection guidance became more outcome-based after the nearly 250 public comments on the draft: it says what must be true, not which hardware to buy.

On the verifier side, the language is unusually direct. Every token must carry an explicit audience, and access control points must reject tokens with a missing or wrong audience and should raise a security alert when they see one. Relying parties must also confirm that the signing key is appropriate for the target environment, which is what stops a key compromised in one tenant from working in another. The IETF’s JSON Web Token Best Current Practices supplies the implementation detail: libraries must let the caller fix the set of accepted algorithms and must not use any other, and should not consume tokens using none unless the caller explicitly asks for it. The none variant of the attack Titan fell to was already documented in that RFC in 2020.

On the holder side, the report asks for short lifetimes, short-lived workload refresh tokens, revocation signals and keeping tokens out of logs and pipelines. Its most consequential line for platform teams is that workload identities should use sender-constrained mechanisms such as mutual TLS or DPoP whenever feasible.

Which controls reach which failure

The practical value of the three-point model is that it shows where a familiar control stops working. The table below is our reading of IR 8587’s threat and mitigation mapping, condensed to the question a platform or security lead actually faces: which of the three failures, forged, blind acceptance or stolen, does each control reach? Sender binding means proof of possession, and verifier checks cover algorithm, audience, issuer and key scope.

ControlReachesMisses
Short lifetimesStolen, by limiting the windowForged, blind
Token revocationStolen, once detectedFresh forgeries, blind
Sender binding (mTLS, DPoP)Stolen, with limits; forged, partlyBlind
Key isolation and scopingForgedBlind, stolen
Key rotation and withdrawalForged, once the old key is withdrawn; stolen, as a side effectBlind
Verifier checksBlind; forged and stolen, partlyNone fully
Usage logs and alertsAll three, as detection onlyPrevention

Verifier checks only partly stop forgery and theft: they reject a key used outside its scope and a token presented to the wrong service, but not a well-formed forgery or a stolen token used where it belongs. Proof of possession raises the bar for forgery, but an attacker holding the issuer’s signing key can generally bind a forged token to a key of their own. Revoking individual tokens cannot reliably contain an attacker who can mint fresh ones. Against forgery, the lever is withdrawing trust in the key, which also invalidates every legitimate token it signed, stolen ones included. Switching to a new signing key achieves nothing while verifiers still accept the old one.

Proof of possession for workloads is now a platform decision

For service-to-service traffic inside Kubernetes, the default credential has been the service account JWT. The Kubernetes post announcing Pod Certificates is candid about its limits: service account tokens are bearer tokens, so “if you have the token, then you are the identity,” and because a workload hands copies to every peer it authenticates to, each peer can impersonate it. The time, object and audience binding that Kubernetes applies are partial mitigations, and the post says none of them is a complete defense.

With Pod Certificates, the kubelet generates a private key for the pod, obtains an X.509 certificate from a signer and rotates both automatically, and the workload authenticates with mutual TLS, proving it holds the key without handing it over. That is the sender-constrained model IR 8587 recommends for workloads.

Adopting it is still a platform project rather than a flag, and three constraints decide the timing.

  • You supply the signer. The Kubernetes project does not yet ship a Pod Certificate signer in core. A cluster that wants these certificates needs a signer controller, which means operating a certificate authority, its keys and its trust bundles. The post’s own example signer is explicitly not a production solution.
  • Applications must handle rotation. Signers that eventually ship in core will issue certificates valid for at most 24 hours, and other signers are capped at 91 days. The kubelet rewrites files as certificates refresh, and applications have to notice through file watching or polling. A service that reads its certificate once at startup will fail when it expires.
  • Not every peer speaks mTLS. Service account tokens work because almost everything can validate a JWT, including cloud provider federation. Certificates cover in-cluster and mesh traffic well, while calls to external APIs may still need a token, where DPoP, defined in RFC 9449, or certificate-bound tokens under RFC 8705 are the sender-constrained options if the other side supports them.

Proof of possession also has a boundary worth stating plainly. RFC 9449’s security considerations note that if an adversary can run code in the client’s context, DPoP’s guarantees no longer hold: even with a non-exportable key, the attacker can generate valid proofs while the client is online. Binding a token to a key stops replay from elsewhere. It does not protect a workload that is already compromised, which is why it complements host and runtime security rather than replacing them.

For teams already running a service mesh or SPIRE, Pod Certificates offer a native delivery path. For teams with neither, weigh the cost of running a signer against bearer tokens that any peer can replay. Where service accounts reach sensitive data or cross trust boundaries, that usually favors doing the work; where they only call the Kubernetes API, audience-bound short-lived tokens may remain adequate for now.

Where teams usually find their gaps

Four antipatterns recur, each invisible from the other failure points.

The first is a verifier nobody tested for rejection. Integration tests prove that valid tokens are accepted, and few prove that invalid ones are refused. The fix is a small negative test suite run against every token-consuming service: alg set to none, a swapped algorithm, a wrong audience, a wrong issuer, an expired token and a token signed by a key from another environment. Each must be rejected, and the audience failures should produce the alert the report calls for.

The second is one signing key everywhere. When a single key signs tokens for every tenant, every environment and every product, compromising it compromises all of them, and rotating it is so disruptive that nobody does. Scoping shrinks that blast radius, and automated rotation makes the change routine.

The third is long-lived credentials hiding behind short-lived ones. A stolen one-hour access token still expires in an hour, but whoever steals a refresh token that never expires, or a static client secret in a CI variable, can keep obtaining fresh access tokens until that credential is revoked. In Kubernetes, the service account documentation still describes Secret-based tokens that neither expire nor rotate. They are not recommended, but they can still be created by hand, and clusters upgraded from versions before 1.24, which generated them automatically, may still carry some.

The fourth is tokens in the paper trail. The report requires that tokens never appear in logs, pipeline output or build artifacts, and that pipelines fetch secrets from a secrets manager and inject them only at runtime. Build logs are widely readable and long-retained, which turns any long-lived token in them into a standing breach.

Where to start, and in what order

A team that wants to act on this does not need to adopt the whole report at once. The order below follows the failure points from cheapest to most structural.

Start with the verifiers, because the fix is local and the exposure is total. Pin each token consumer’s accepted algorithms, require audience and issuer checks, and add the negative test suite to its pipeline. The work is local, and it closes a failure no other investment can.

Then look at the issuer. Find out how many signing keys exist, what each can sign for, where it is stored and when it last rotated. Split keys by environment and, where possible, by tenant, and rehearse a rollover before you need one.

Then shorten the holder’s exposure. Bring access tokens to an hour or less, shorten workload refresh tokens, remove Secret-based service account tokens, scrub tokens from logs, and subscribe relying parties to revocation signals where available.

Finally, choose where proof of possession pays for itself. For identities that reach sensitive data or cross trust boundaries, decide between Pod Certificates, an existing mesh or SPIRE, and DPoP or certificate-bound tokens for external APIs, and give the capability a platform owner.

IR 8587’s real contribution is to put issuers, verifiers and holders on the same page. The question to take from it is not how short your tokens live but which of the three failure points your organization has actually tested.

Token security usually falls between the teams that run the identity provider, the platform and the services, which is why its gaps survive audits. We help engineering and security teams test every token consumer for rejection, design signing key scoping and rotation, and plan workload identity with mTLS or proof-of-possession tokens where the risk justifies it. Talk with us about your token and workload identity design.

Security Identity Architecture Kubernetes
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.