Every Hop Needs Its Own Token
How to carry a user identity through an MCP server with OAuth On-Behalf-Of flow.
A repo got deleted. You pull the audit log, and every row says mcp-service-account. Not Alice. Not Bob. Somebody’s AI agent did it, and there’s no record of whose.
It starts somewhere completely reasonable.
MCP servers are quickly becoming the interface between AI agents and the tools we use every day. GitHub, Slack, Microsoft 365, internal APIs, data platforms, infrastructure systems, etc, all of these systems are becoming reachable through MCP. Users like this because they can read and change real things without leaving their AI assistants.
The part that usually gets skipped is identity. Plenty of MCP servers are left open, or protected by a single service token, or authenticated in a way that loses track of the human as soon as the request moves one hop further. That’s fine for a demo but not in production.
So if an AI agent reads a document, queries customer data, or restarts a service, you need answers to three questions:
- Who asked for this?
- Were they allowed to do it?
- Can you prove both later?
Authenticating at the MCP endpoint only answers them for the first hop. The identity has to survive the next one too. This post is about how OAuth’s On-Behalf-Of flow carries it from an MCP server to a downstream API, without reusing the original access token.
The shared-token trap
Say a company has an internal project-management API: projects, tickets, comments, team assignments, like Jira. They want a few of those operations exposed through an MCP server so employees can work tickets from their assistant.
It usually gets set up like this:
Alice ─┐
Bob ─┼──> MCP server ──> Project API as a service account
Carol ─┘
The project API doesn’t see Alice, Bob, or Carol anymore. It sees the MCP service account. And that service account is usually a lot more powerful than any of them, so now:
- Every user inherits the full reach of the service account, which almost always has broader access than any single person.
- Removing a user doesn’t remove their effective access through the tool.
- Downstream audit logs lose the person responsible for the action.
The other common version is handing every employee a personal API token to paste into their MCP client. That works for three people, gets annoying at 30, and is unmanageable at 3000. Someone has to distribute, store, rotate, revoke, and map all of them, and the mapping changes every time somebody switches teams.
We already solved this for normal apps with single sign-on. So why not just use the same identity for MCP?
SSO only covers the first hop
You can, and you should. The user signs in before the MCP client connects to a remote MCP server. The client opens an authorization flow, the user authenticates with the company identity provider, and the client gets back an access token meant for the MCP server.
Then it can make an authenticated tool call:
POST /mcp HTTP/1.1
Host: mcp.example.com
Authorization: Bearer <token-a>
The token carries claims along these lines:
{
"iss": "https://login.example.com/tenant-id",
"aud": "api://enterprise-mcp",
"scp": "mcp.invoke",
"oid": "user-123",
"tid": "tenant-456"
}
oid and tid identify the user and tenant. scp is the delegated permission. The claim that matters most here is aud, the audience: the service this token was issued for.
The MCP server accepts it, because that’s who it was addressed to. The project API won’t. It expects a token for project-api, and this one says enterprise-mcp. That rejection is the point. The whole job of aud is to stop a token from being reused somewhere it wasn’t issued for.
Access tokens are addressed envelopes
Think of an access token as a sealed envelope with a name on the front. The identity provider seals it and writes the recipient. The claims inside describe who the subject is and what they were granted.
An envelope addressed to the MCP server should only be opened by the MCP server. A downstream API should reject it even though the same identity provider signed it, and even though there’s a perfectly real user identity sitting inside.
The shortcut everyone reaches for is forwarding the MCP token as-is:
AI client ── Token A ──> MCP server ── Token A ──> Downstream API
That’s token passthrough, and the MCP authorization specification forbids it: “The MCP server MUST NOT pass through the token it received from the MCP client.”
So the user’s identity should cross service boundaries but the original access token shouldn’t. Every service should get its own.
Every hop needs its own token
This is what OAuth delegation is for. The MCP server takes the user token that was issued to itself, and asks the identity provider for a second token issued to the downstream API. In Microsoft Entra that’s the OAuth 2.0 On-Behalf-Of (OBO) flow.
MCP client
│ Token A (aud: MCP server)
v
MCP server
│ Token A + workload credential
v
Identity provider
│ Token B (aud: Project API)
v
MCP server
│ Token B
v
Project API
What’s happening here is that two identities go into the exchange: the user, from the incoming access token, and the MCP server itself, from its own workload credential. That second identity is the part we usually forget, and it’s what lets the hop from the MCP server to the project API carry the user along with it.
You come out with two tokens:
Token A
Audience: MCP server
Permission: mcp.invoke
User: Alice
Token B
Audience: Project API
Permission: access_as_user
User: Alice
Different envelopes, same Alice inside. All Token B says is that this call is being made as Alice. The project API takes it from there and applies Alice’s actual team and project permissions, instead of the permissions of one very generous service account.
The flow, step by step
Say the MCP server has a list_my_tickets tool that calls GET /v1/tickets?assignee=me on the internal project API.
1. The user signs in through the MCP client
The client runs an authorization-code flow with PKCE. The user authenticates with the enterprise identity provider and grants the access being asked for. Back comes Token A, with the MCP server as its audience.
2. The MCP client calls the tool
Token A goes in the authorization header. Before running the tool, the server validates at least:
- The signature.
- The issuer and tenant.
- The audience.
- Expiration and not-before.
- The required delegated scope.
- The calling client, if you restrict clients.
3. The MCP server asks for a project API token
The MCP server then sends an OBO request to the identity provider with Token A as the user assertion, its own client identity, and the delegated scope it needs on the project API.
In Entra, that request looks like this:
POST /<tenant>/oauth2/v2.0/token HTTP/1.1
Host: login.microsoftonline.com
Content-Type: application/x-www-form-urlencoded
client_id=<mcp-server-client-id>
&grant_type=urn:ietf:params:oauth:grant-type:jwt-bearer
&assertion=<token-a>
&requested_token_use=on_behalf_of
&scope=api://project-api/access_as_user
&client_assertion_type=urn:ietf:params:oauth:client-assertion-type:jwt-bearer
&client_assertion=<mcp-workload-assertion>
The exact shape depends on your identity provider, library, and credential type. Most other providers do this through the standard token exchange grant instead, which I get into further down.
4. The identity provider issues Token B
The IDP checks that:
- Token A is valid and was issued for the MCP server.
- The MCP server successfully authenticated itself.
- The MCP application is allowed to request that delegated project API scope.
- The user or an administrator has granted consent.
If all of that passes, you get Token B, addressed to the project API.
5. The MCP server calls the project API
GET /v1/tickets?assignee=me HTTP/1.1
Host: projects.example.com
Authorization: Bearer <token-b>
The project API validates Token B and applies its own authorization rules. Alice might read tickets in one project and get a clean 403 on another, exactly like they would in the web UI. And Token B never travels back to the MCP client.
What you have to register
The names differ between identity providers, but the pieces are usually the same.
The MCP API. The MCP server is registered as a protected API that exposes a delegated scope, something like api://enterprise-mcp/mcp.invoke. The MCP client requests that scope, and the token it gets back is meant for the MCP server.
The downstream API permission. The project API is registered as its own protected resource, and it exposes a delegated scope the MCP application is granted:
api://project-api/access_as_user
All that scope says is that the MCP server is allowed to call the project API as the signed-in user. What Alice can actually do is decided inside the project API, from Alice’s identity and the permission model the API already has. That’s usually where you want that decision, because the project API is the thing that knows about teams, projects, and roles. Copying all of that into scopes means keeping two permission models in sync forever.
RFC 8693 and Entra OBO aren’t the same thing
RFC 8693 defines the standard OAuth 2.0 Token Exchange grant:
urn:ietf:params:oauth:grant-type:token-exchange
Entra OBO uses a JWT bearer grant with an extra parameter:
grant_type=urn:ietf:params:oauth:grant-type:jwt-bearer
requested_token_use=on_behalf_of
Same architectural idea, different wire protocol. Keycloak, Auth0, Ping, and others now implement the 8693 grant. Entra’s OBO shipped before RFC 8693 was published, which is most of why it looks different.
Worth knowing: 8693 defines the mechanism but leaves the policy to the provider. Which subject tokens are valid, which audiences a client may ask for, whether you get delegation or impersonation, all of that is up to the implementation. So two providers can both support 8693 and still not behave the same way. What you can use depends on your identity provider and the downstream service. The name on the box matters less than what has to stay true either way: the downstream service gets its own token, addressed to it, carrying a user context somebody explicitly delegated.
Third-party SaaS is a different problem
One thing to be clear about: this whole post assumes one organization owns both APIs. That’s the cleanest OBO case, because both sides can point at the same identity provider and everything lines up.
A vendor’s API is its own thing. Most of them only accept tokens they issued, so nothing you set up on your side gets you a Token B for their audience. That’s a separate post. If you’re there now, Cross-App Access and MCP’s Enterprise-Managed Authorization extension are where to start reading.
The security baseline
OBO doesn’t make an MCP server secure on its own. At minimum:
- Validate the issuer, tenant, audience, signature, expiration, and scopes on every incoming token.
- Reject tokens issued for any resource other than the MCP server.
- Never pass the incoming MCP token to a downstream API.
- Request only the downstream scopes the MCP server needs.
- Keep the real permission check in the downstream API.
That’s the floor. Dynamic client registration, OAuth proxying, consent confusion, token substitution, tenant mix-ups, bearer-token replay, over-permissioned tools, and prompt injection each open their own attack paths.
MCP adds one more service boundary to the identity architecture you already have, and it behaves like every other one. The token presented to the MCP server belongs to the MCP server. The token presented to the project API belongs to the project API. Get that right and the next incident review is a much shorter meeting, because the log says Alice.
The identity should survive the tool call. The token shouldn’t.