Guide

Building a Coding Agent Sandbox Around GitHub

A sandbox setting protects your filesystem. It does nothing for the GitHub credential your agent carries or the workflow runner that executes what it pushes.

Quick Answer

An isolated system environment for a coding agent that works against GitHub needs three separate boundaries, not one sandbox setting. The first is where the agent's commands actually run: a container or VM such as a GitHub Codespace, defined by a devcontainer.json file, or a sandbox on your own machine. The second is the GitHub credential it authenticates with: a fine-grained personal access token scoped to named repositories and permissions, or a GitHub App installation token, rather than a classic token that reaches your whole account. The third is the automation boundary around anything the agent triggers in GitHub Actions: a GITHUB_TOKEN restricted to read access by default, and runner groups that keep self-hosted runners away from public repositories. Treat these as three separate decisions, choose each on purpose, and the result is a real boundary instead of one switch you are hoping covers everything.

Why one sandbox setting is not enough

A coding agent that works against GitHub touches three different systems at once: the machine its commands run on, the GitHub identity it authenticates as, and whatever Actions workflow picks up what it pushes. A sandbox that only limits the filesystem leaves the other two wide open.

This is why a search like 'coding agent sandbox github' rarely has one clean answer. The phrase bundles three separate engineering decisions, each with its own tools and its own way of going wrong.

Get the execution environment right and you can still hand the agent a personal access token that reaches every repository you own. Get the token right and a workflow it triggers can still run with far more access than the task needed. The three layers have to be solved separately, because fixing one does nothing for the other two.

  • Execution: where the agent's commands actually run.
  • Credential: what identity it authenticates as on GitHub.
  • Automation: what a workflow it triggers is allowed to do.

The execution environment: GitHub Codespaces and devcontainer.json

A GitHub Codespace runs your project inside a development container hosted on a virtual machine, configured by a devcontainer.json file. That file is not a GitHub-only format. It is a vendor-neutral specification maintained by the Dev Containers project, so the same configuration also runs locally through an editor and Docker, with or without a Codespaces subscription at all.

Codespaces secrets look like a vault but are not one. GitHub exports them as plain environment variables once the codespace starts, capped at 100 secrets per account and 48 KB each, and they are not available during the container's own build step. Anything running inside that terminal session, including an agent, can read them the same way it would read a .env file on a laptop.

So a Codespace isolates where the agent's commands run. On its own, it does nothing to protect what the agent can read once it is running there, which is a separate problem with a separate fix.

  • devcontainer.json works locally too, not only on GitHub's cloud VMs.
  • Codespaces secrets become environment variables, visible to anything in the session.
  • Secrets are unavailable at build time, only after the container starts.

The GitHub credential: fine-grained tokens over classic ones

A classic personal access token reaches every repository in every organization you belong to, plus your personal account. GitHub's own documentation now recommends fine-grained personal access tokens instead, because they can be limited to one organization and to the specific repositories you name, with named permissions rather than broad scopes.

A GitHub App goes a step further. It authenticates with its own key to request an installation access token, scoped to only the repositories the app was installed on, and the resulting activity shows up as the app rather than as your personal account. That separation matters when the thing requesting access is an automated agent and not you at a keyboard.

Whichever you pick, the underlying rule stays the same: hand the agent a credential that can do the one job in front of it, not a credential that can do everything you personally can.

  • Classic tokens: broad, account-wide access.
  • Fine-grained tokens: limited to named repositories and named permissions.
  • GitHub App tokens: scoped to the app's installation, not your account.

The automation boundary: GITHUB_TOKEN and self-hosted runners

If the agent's work ends in a GitHub Actions workflow, that workflow gets its own automatic GITHUB_TOKEN. GitHub's security guidance recommends granting it the least access it needs, and setting the repository or organization default to read-only on contents, so an individual workflow has to ask explicitly for anything more than that.

Self-hosted runners raise the stakes further. GitHub is direct about this in its own hardening guide: a self-hosted runner should almost never serve a public repository, because any user can open a pull request against that repository and run code on the runner behind it. Runner groups let you keep a runner scoped to the specific repositories and organizations allowed to use it.

None of this is specific to AI agents. It is the same hardening GitHub asks of any automated contributor, which is exactly the point: an agent pushing code is an automated contributor, and it should be governed the same way.

  • Default GITHUB_TOKEN permissions to read-only, then widen per job.
  • Keep self-hosted runners off public repositories entirely.
  • Use runner groups to scope which repositories can use which runner.

Putting the three boundaries together

A setup that actually holds picks one answer for each layer. Run the agent's commands inside a container, whether that is a Codespace, a local devcontainer, or a sandbox you built yourself. Authenticate it with a fine-grained token or a GitHub App scoped to only the repositories the task needs. Default any workflow it can trigger to the least access GitHub will let you set.

Each layer stays replaceable on its own once it is separated out this way. You can swap the execution environment without touching the credential, and rotate the credential without touching the runner policy. That is what makes the setup something you can maintain, rather than a pile of settings nobody wants to revisit after the first week.

  • Pick one container or sandbox strategy and use it consistently.
  • Scope the GitHub credential before worrying about anything else.
  • Review the three layers separately whenever something changes.

Where Forkbench fits, and where it does not

Forkbench runs coding agents in real terminals on your own Mac, not in a hosted container somewhere else. It can lock a Thread to the folders it is allowed to touch, enforced by the macOS kernel sandbox, which covers the first layer above for work you keep on your own machine. Its Vault lets a command use a GitHub token by name, so the value never reaches the prompt, the command line, or the transcript.

What it does not do is reach into GitHub's own infrastructure. Codespaces, devcontainer builds, and Actions runners are GitHub's boundaries, not Forkbench's, and nothing here implies the two talk to each other. The folder lock is also opt-in and does not restrict the network, so a locked agent can still send out anything it is allowed to read.

  • Folder lock: opt-in, local to the Mac, enforced by the kernel sandbox.
  • Vault: a GitHub token is used by name, never shown to the agent.
  • Not covered: Codespaces, Actions runners, or anything past your own machine.

A setup order that holds up

Start with the credential, since it is the cheapest piece to fix and the most damaging to leave wrong. Move to the execution environment once the credential is scoped down. Add the Actions-side defaults last, because they only matter once the agent's work actually reaches a workflow run.

Revisit all three the next time something changes: a new repository, a new runner, or a new agent joining the project. A boundary that was correct in January can be quietly wrong by October if nobody checks it again.

  • Replace any classic PAT the agent uses with a fine-grained token or a GitHub App.
  • Choose a container or sandbox for where its commands run, and stick with it.
  • Set the repository's default GITHUB_TOKEN permission to read-only.
  • If you use self-hosted runners, confirm none of them serve a public repository.
  • Keep the token out of the project folder itself, in a vault or the Keychain.

Related: Give an agent deploy access without the credential, macOS Seatbelt and coding agents, Sandbox AI coding agents on macOS safely, Download Forkbench

Frequently asked

  • Does GitHub Codespaces isolate my agent from my repository's secrets?

    No. Codespaces secrets are exported as plain environment variables once the codespace starts, so anything running in that session, including an agent, can read them the same way it would read a .env file.

  • What is the difference between a classic and a fine-grained GitHub token?

    A classic token reaches every repository you have access to across every organization. A fine-grained token is limited to the organization and repositories you choose, with named permissions, and GitHub now recommends it as the default.

  • Should a coding agent be allowed to use a self-hosted Actions runner on a public repository?

    No. GitHub's own hardening guidance says a self-hosted runner should almost never serve a public repository, because anyone can open a pull request that runs code on that runner.

  • Does Forkbench integrate with GitHub Codespaces or Actions?

    No. Forkbench locks folders and manages secrets on your own Mac. Codespaces, devcontainer builds, and Actions runners are separate systems that GitHub operates on its own infrastructure.

  • Can devcontainer.json be used without GitHub Codespaces?

    Yes. It is a vendor-neutral specification, so the same file also configures a local container through an editor and Docker, with no Codespaces subscription needed.

Keep reading