Guillaume Duvernay

Your MCP can't accept binary files, here's the fix

MCPagentssecurityAI Glot

My SaaS AI Glot translates files (CSV, XLSX, Word, PDF and a few dozen other formats) and it has an MCP server. The first version had a gap I knew about: the agent could send me a public link to a file, or a small piece of text, but a file sitting on your computer had no way in.

I fixed it last week. The trick is simple, a bit hacky, and I think it works for any MCP that needs to receive a file, so here it is.

Why a file doesn’t fit in a tool call

The arguments of an MCP tool are JSON. To put a binary file in there, you have to encode it as base64, which is text, and the model has to write that text itself, character by character, inside the tool call.

The 2.7 million is base64 itself: four characters for every three bytes. I didn't convert it to tokens, since that depends on the model.

It’s the same problem as the 200 CSV rows in why I stopped using MCP for most of my work, except here one call is already the whole job.

And that’s if the agent can even read your file. Connecting an MCP server doesn’t give an agent access to your attachments or your disk. That comes from the app around it.

The trick: the MCP gives a URL, the agent sends the file there

The agent doesn’t have to send the file through MCP. It only needs to know where to send it. So the MCP returns an upload URL, and the agent does the transfer itself with a normal HTTP request.

The model only writes small JSON. The file goes from your machine straight to storage and the model never reads it.

In AI Glot the first tool is create_upload_url. You give it the filename and the exact size in bytes:

{ "filename": "catalogue.xlsx", "size_bytes": 15961 }

It returns an upload_id, a URL, the method (PUT), temporary headers, two deadlines, and instructions written for the agent. Then the agent runs something like:

curl -X PUT "$UPLOAD_URL" \
  -H "Authorization: $TEMP_UPLOAD_TOKEN" \
  -H "Content-Type: application/octet-stream" \
  --data-binary @catalogue.xlsx

After that it calls create_translation with only the upload_id. No filename, no content, no base64. It gets a plan, asks you to confirm, and approve_translation spends the credits.

I tested it live on release day with a 15,961-byte XLSX containing two words. The PUT returned 200, the plan counted two words, and after approval the translation cost one credit. Creating the upload and the plan cost nothing.

The other route, and why it’s not enough

AI Glot already accepted a file_url: you give a link, and my server downloads the file itself. That works well when the file is already online.

But most of the time the file is on your computer. To use file_url you’d first have to upload it somewhere, make it publicly accessible, and paste the link. That’s a lot of steps to ask someone who just wants to say “translate this spreadsheet”. Small text files can still go in directly as content, and that’s it.

So there are now three ways in: text content, a public URL, and the upload URL for the file that only exists on your machine.

When the agent can’t send a PUT

This is the limit. The agent needs to read your file and send an HTTP request. Claude Code or Codex can, because they have a shell. A chat window without a code sandbox can’t, even if the MCP is connected.

So the MCP has to say it, everywhere the agent will read: the tool description, the tool result, the API docs and llms.txt. Mine tell the agent that if it can’t send the attachment, it should tell the user clearly and suggest the other ways: the web app, the CLI, or a public link. They also tell it never to say a file was uploaded when it wasn’t, and never to approve a payment without a verified file and plan.

Without it, an agent that can’t send the PUT just fails quietly, or worse, pretends it worked.

How to secure it

You’re letting an unknown HTTP client write into your storage, so this is the part where I spent the most time. The idea is that the secret you return can do one thing and nothing else.

The agent holds both and they never swap. The file travels with the weak one, the instructions travel with the strong one.

What I did, and what I’d do again on any server:

  • The secret is in a header, never in the URL. URLs end up in logs and history. The URL alone gives access to nothing.
  • It only works for one session. It allows a PUT on that session’s endpoint and nothing else. I checked: sent to the account API, it returns a 401.
  • I store a hash, not the secret. It’s an HMAC salted with the session ID. The plaintext is returned once with Cache-Control: no-store, and my logs only contain IDs and statuses.
  • One attempt per session. A second PUT returns a 409 and can’t replace the file. If the first one fails, the agent asks for a new session.
  • The size is declared, then enforced. The server rejects unsupported extensions and oversized files before issuing anything, counts every byte it receives, refuses compressed or short or extra bodies, and gives up after 60 seconds.
  • Everything expires. The secret after 10 minutes, the file after one hour if no translation uses it, and a daily job deletes what’s left after 24 hours.
  • The caller doesn’t choose where it lands. No filename, no path, no workspace in the request. I generate the storage key inside that workspace’s own prefix.
  • Only the same credential can use the upload. If a different connection or workspace tries to create a translation from it, it’s refused.
  • Revoking works. If someone removes a connection or its permission, its unused upload secrets stop working.
  • Quotas. 20 pending uploads per workspace and 120 per hour, on top of the normal rate limits.
  • Uploading costs nothing. Only approval spends credits, so nobody can drain a balance with uploads.

I didn’t use a native presigned storage URL because those can be reused until they expire, and I wanted exactly one attempt with a size limit. So there’s a small Worker in front of the bucket that keeps that state in the database.

Each arrow is a compare-and-set in the database, so two requests can't both win the same session, and nothing goes back to an earlier state.

Lost responses

Most of the bugs I had to think about were requests that worked while the caller never found out.

If the agent loses the response to its PUT, it doesn’t upload again. It tries create_translation with the same upload_id first. If the file arrived, that works. If not, it asks for a new session. Overwriting an old one isn’t possible.

Creating the translation is idempotent too: the upload_id is the key, so repeating the call returns the same translation. I repeated it in my test and got exactly one. A repeat also never re-plans with a changed instruction. For that there’s a separate tool.

And when the server doesn’t know what happened, it never deletes the file. If a database commit might have gone through, the bytes stay and the cleanup job sorts it out later.

If you want to copy it

The minimum I’d build:

  1. A tool that returns an upload URL and the headers to use, so the model never sees the bytes.
  2. A secret that is single-use, short-lived, tied to one session and kept out of the URL.
  3. A size the server enforces, and a storage path the server picks.
  4. Instructions in the tool result for agents that can’t send a PUT.
  5. Creation keyed on the upload ID, so a retry is safe.

The MCP side is a few lines. Most of the work was the failure cases.

Sources