Isolation You Can't Prove Is Just an Assumption
The environment bugs that hurt do not throw.
A variable that should point at staging points at production, and every request succeeds. A service account created for development gets a role on the production project, granted in a hurry to unblock someone, and nothing breaks. A new secret gets added to three environments out of four. Each of these is invisible right up to the moment something writes where it should not, and then it is an incident with a very short and very embarrassing root cause.
Almost every team I have worked with documents its environment isolation. There is a diagram with three boxes, a wiki page with a naming convention, maybe a checklist in the release process. What almost nobody has is a check that fails when the isolation stops being true.
That is the same distinction this series keeps coming back to. A boundary you only describe is a wish. A boundary you enforce is architecture. And an enforced boundary you never test is an assumption that happens to be holding today.
Here is the ladder I use, four levels, each with the code to do it. On a recent platform I built the first two properly. The other two are what I would add now, and I will be explicit about which is which, because the gap between them is the useful part.
Level 1: refuse to boot with the wrong configuration
The cheapest level, and the one that catches the most. Every deployable, frontend and backend, validates its configuration at startup against a schema. A missing key, a value of the wrong shape, a URL that is not a URL: the process does not start, and it says exactly why.
I use zod for this, because the schema doubles as documentation. Every property carries a description, and the description ends up in the error message, so the person reading the failure at two in the morning learns what the key is for, not only that it is missing.
// src/env.schema.ts
import { z } from 'zod'
export const EnvObject = z.object({
APP_ENV: z
.enum(['dev', 'staging', 'prod'])
.describe('The environment this process believes it is running in'),
GCP_PROJECT_ID: z
.string()
.regex(/^acme-app-(dev|staging|prod)$/)
.describe('The Google Cloud project hosting this deployable'),
DATABASE_URL: z
.string()
.url()
.describe('Postgres connection string, injected from Secret Manager'),
PAYMENTS_API_KEY: z
.string()
.min(20)
.describe('Payments provider key, injected from Secret Manager'),
LOG_LEVEL: z
.enum(['debug', 'info', 'warn', 'error'])
.default('info')
.describe('Minimum log level'),
})
// The process must agree with its own project about where it is.
export const EnvSchema = EnvObject.refine(
env => env.GCP_PROJECT_ID.endsWith(`-${env.APP_ENV}`),
{
path: ['GCP_PROJECT_ID'],
message: 'does not match APP_ENV',
}
)
// Keys that live in Secret Manager. By convention the secret id equals the key.
export const SECRET_KEYS = ['DATABASE_URL', 'PAYMENTS_API_KEY'] as const
export const REQUIRED_KEYS = Object.entries(EnvObject.shape)
.filter(([, schema]) => !schema.isOptional())
.map(([key]) => key)
// src/env.ts
import { EnvObject, EnvSchema } from './env.schema'
export function loadEnv(source: NodeJS.ProcessEnv = process.env) {
const result = EnvSchema.safeParse(source)
if (result.success) return result.data
const shape = EnvObject.shape as Record<string, { description?: string }>
const lines = result.error.issues.map(issue => {
const key = issue.path.join('.')
const description = shape[key]?.description
return ` ${key}: ${issue.message}${description ? ` (${description})` : ''}`
})
console.error(`Refusing to start, invalid configuration:\n${lines.join('\n')}`)
process.exit(1)
}
The cross-check at the bottom of the schema is the part that earns its place in a post about isolation rather than validation. The process knows which environment it thinks it is in, and it refuses to run if its own project disagrees. A development build accidentally deployed with production configuration does not quietly serve traffic. It dies on the first line.
On a platform that will not route traffic to a revision that fails to start, which is the case on Cloud Run and on anything with a readiness probe, this is already a deployment gate, and it cost nothing extra.
Its one limit is timing. It fails at boot, which means after the deploy. Level 4 moves the same check in front of the deploy, using the same schema, so hold that thought.
Level 2: make the cross-reference impossible to express
The obvious next step is a lint that scans development configuration for anything that looks like production: a production hostname, a production project id, a secret name with prod in it. It works, and it is a detector, which means it only catches the patterns somebody thought to write down.
What I did instead was remove the possibility. Every environment is its own Google Cloud project, provisioned by the same Terraform module with a different variable file. Its own secrets, its own service accounts, its own infrastructure, no shared anything.
# environments/prod/main.tf
module "platform" {
source = "../../modules/platform"
project_id = "acme-app-prod"
environment = "prod"
secret_ids = ["DATABASE_URL", "PAYMENTS_API_KEY"]
}
# modules/platform/identity.tf
resource "google_service_account" "runtime" {
project = var.project_id
account_id = "runtime"
display_name = "Runtime identity (${var.environment})"
}
resource "google_secret_manager_secret" "this" {
for_each = toset(var.secret_ids)
project = var.project_id
secret_id = each.value
replication {
auto {}
}
}
resource "google_secret_manager_secret_iam_member" "runtime_access" {
for_each = google_secret_manager_secret.this
project = var.project_id
secret_id = each.value.secret_id
role = "roles/secretmanager.secretAccessor"
member = "serviceAccount:${google_service_account.runtime.email}"
}
The effect is that there is nothing for a lint to find. A development service resolves DATABASE_URL in the development project because that is the only project it knows about. To reach a production secret it would have to name another project explicitly, and even then its identity has no permission there.
This is the move I would make first on any new platform, and it is worth saying why it beats the lint. A detector has to anticipate every way the mistake can be written. A structure where the mistake cannot be written does not need to anticipate anything.
What separation does not give you
Here is where I stopped, and where I should not have.
Separate projects are a state, and states drift. The structure makes the mistake impossible to express by accident. It does not stop someone expressing it on purpose. An urgent fix, a migration script that needs to read production data from a development runner, a colleague who grants the development service account a role on the production project "just for today". One binding, added through the console, outside Terraform, and the isolation is gone without a single test turning red.
That is the gap between levels 2 and 3. Everything so far makes isolation the default. Nothing so far tells you when the default has been overridden.
Level 3: prove the other environment says no
The strongest evidence that two environments are isolated is to try to cross the boundary and watch it refuse. So the test is negative: impersonate the development runtime identity and attempt to read a production secret. The test passes only on a permission error.
#!/usr/bin/env bash
# scripts/check-cross-env-denied.sh
# Passes only if the dev runtime identity is DENIED access to a prod secret.
set -uo pipefail
DEV_SA="runtime@acme-app-dev.iam.gserviceaccount.com"
PROD_PROJECT="acme-app-prod"
PROBE_SECRET="DATABASE_URL"
ERR="$(mktemp)"
if gcloud secrets versions access latest \
--secret="$PROBE_SECRET" \
--project="$PROD_PROJECT" \
--impersonate-service-account="$DEV_SA" >/dev/null 2>"$ERR"; then
echo "FAIL: $DEV_SA can read $PROBE_SECRET in $PROD_PROJECT" >&2
exit 1
fi
if ! grep -q "PERMISSION_DENIED" "$ERR"; then
echo "INCONCLUSIVE: expected PERMISSION_DENIED, got:" >&2
cat "$ERR" >&2
exit 1
fi
echo "OK: cross-environment read denied"
The second if is the important one. A negative test that passes for the wrong reason is worse than no test, because it hands you confidence. If the command fails because the CI identity is not allowed to impersonate the development account, or because of a network error, or because someone renamed the secret, you have not proved isolation. You have proved that something went wrong. So anything other than an explicit permission denial fails the check. The CI identity needs roles/iam.serviceAccountTokenCreator on the development service account for the impersonation to work at all.
The behavioural test covers one secret and one identity. The structural complement asserts that no service account from another project is bound anywhere in the production project's IAM policy.
#!/usr/bin/env bash
# scripts/check-no-foreign-iam.sh
# Fails if service accounts from other projects are bound in prod.
set -euo pipefail
PROD="acme-app-prod"
members="$(gcloud projects get-iam-policy "$PROD" --format=json \
| jq -r '.bindings[].members[]')"
# User-managed service accounts look like name@PROJECT.iam.gserviceaccount.com.
# Google-managed service agents live under gcp-sa-* and are expected.
foreign="$(printf '%s\n' "$members" \
| grep -E '^serviceAccount:.+@.+\.iam\.gserviceaccount\.com$' \
| grep -vE "@(${PROD}|gcp-sa-[a-z0-9-]+)\.iam\.gserviceaccount\.com$" \
|| true)"
if [[ -n "$foreign" ]]; then
echo "FAIL: service accounts from other projects are bound in $PROD:" >&2
echo "$foreign" >&2
exit 1
fi
echo "OK: no foreign service accounts bound in $PROD"
Notice set -euo pipefail at the top, and the || true exactly where an empty result is legitimate. Without them, a failed gcloud call produces an empty member list, the grep finds nothing foreign, and the check reports success. The check itself has to fail closed, or it is just another place where isolation can silently stop being true.
Both scripts verify. If you want the platform to enforce the boundary even when a bad binding gets through, the native tool on Google Cloud is VPC Service Controls: a perimeter around the production project that blocks access to its data from outside, whatever IAM says.
resource "google_access_context_manager_service_perimeter" "prod" {
parent = "accessPolicies/${var.access_policy_id}"
name = "accessPolicies/${var.access_policy_id}/servicePerimeters/prod"
title = "prod"
status {
resources = ["projects/${var.prod_project_number}"]
restricted_services = [
"secretmanager.googleapis.com",
"storage.googleapis.com",
"sqladmin.googleapis.com",
]
}
}
It needs an organisation and an access policy, and it has a learning curve, so it is not where I would start. For a platform holding regulated data it is where I would end up, because it moves the guarantee underneath the identity instead of trusting the identity to be configured correctly.
Level 4: gate the promotion
Back to the limit of level 1. The schema fails at boot, after the deploy. The same schema can fail before it.
The schema already knows which keys are required and which of them live in Secret Manager. So before promoting a build into an environment, a script checks that every required key is declared on the target service, and that every secret-backed key has an enabled version in the target project. If anything is missing, the promotion stops, with a list, before anything is deployed.
// scripts/check-promotion.ts
// Usage: npx tsx scripts/check-promotion.ts <project> <region> <service>
import { execFileSync } from 'node:child_process'
import { REQUIRED_KEYS, SECRET_KEYS } from '../src/env.schema'
const [project, region, service] = process.argv.slice(2)
if (!project || !region || !service) {
console.error('Usage: check-promotion <project> <region> <service>')
process.exit(2)
}
const gcloud = (...args: string[]) =>
execFileSync('gcloud', [...args, `--project=${project}`], {
encoding: 'utf8',
}).trim()
// Keys actually declared on the target Cloud Run service.
const spec = JSON.parse(
gcloud('run', 'services', 'describe', service, `--region=${region}`, '--format=json')
)
const declared = new Set<string>(
(spec.spec.template.spec.containers[0].env ?? []).map(
(e: { name: string }) => e.name
)
)
const hasEnabledVersion = (secret: string) => {
try {
return (
gcloud(
'secrets', 'versions', 'list', secret,
'--filter=state=ENABLED', '--limit=1', '--format=value(name)'
) !== ''
)
} catch {
return false // missing secret or no access: treat as unresolvable
}
}
const problems = [
...REQUIRED_KEYS.filter(key => !declared.has(key)).map(
key => `${key}: not declared on ${service}`
),
...SECRET_KEYS.filter(key => !hasEnabledVersion(key)).map(
key => `${key}: no enabled version in Secret Manager`
),
]
if (problems.length > 0) {
console.error(`Promotion to ${project} blocked:\n ${problems.join('\n ')}`)
process.exit(1)
}
console.log(`OK: every required key resolves in ${project}`)
Every error path ends in a blocked promotion. A secret that does not exist, a CI identity that cannot list versions, a service that cannot be described: none of them is read as success.
Wired into the pipeline, the promotion job runs all three checks before the deploy step, so a failure anywhere stops the release.
# .github/workflows/deploy.yml (excerpt)
promote-prod:
needs: [build, test]
runs-on: ubuntu-latest
environment: prod
permissions:
contents: read
id-token: write
steps:
- uses: actions/checkout@v4
- uses: google-github-actions/auth@v2
with:
workload_identity_provider: ${{ vars.WIF_PROVIDER }}
service_account: ${{ vars.CI_SERVICE_ACCOUNT }}
- uses: google-github-actions/setup-gcloud@v2
- run: npm ci
- run: npx tsx scripts/check-promotion.ts acme-app-prod europe-west1 api
- run: ./scripts/check-no-foreign-iam.sh
- run: ./scripts/check-cross-env-denied.sh
- run: ./scripts/deploy.sh prod
The schema is now the single source of truth in three places: it refuses to boot a misconfigured process, it blocks a promotion into an environment that cannot satisfy it, and it documents every key for whoever reads the failure.
The order I'd do it in
If I were starting a platform tomorrow, I would not climb the ladder one rung at a time.
Structural separation first, because it removes a whole category of mistake instead of detecting instances of it. The startup schema the same day, because it is an afternoon of work and it turns every configuration error from silent into loud. Then the two things I skipped last time: the negative test and the foreign-binding check, in CI and on a schedule, because separation holds only until somebody overrides it, and the override is always urgent and always reasonable. The promotion gate last, once there is more than one environment worth protecting.
And every check fails closed, including the checks themselves. An isolation test that reports success when it could not run is the most dangerous thing in the pipeline, because it is the one everybody trusts.
So the question I would ask of your own setup is a short one. If someone granted your development identity a role on production this afternoon, which of your checks would turn red, and when?