Configuration
gc-batch uses sensible defaults but allows you to override settings via command-line options and environment variables.
Default Settings
| Setting | Default | Description |
|---|---|---|
| Location | us-central1 |
GCP region for Batch jobs |
| Project ID | none | GCP project, from --project-id, $GOOGLE_PROJECT, or the default_project_id setting. Required. |
| Machine Type | e2-standard-2 |
VM machine type |
| Boot Disk Size | 30 GB |
Size of the boot disk |
| Boot Disk Type | pd-balanced |
Type of boot disk |
| Input Directory | /mnt/input |
Container mount point for input |
| Output Directory | /mnt/output |
Container mount point for output |
Environment Variables
GCP Configuration
| Variable | Description |
|---|---|
GOOGLE_PROJECT |
Default GCP project ID (used if --project-id not specified) |
USER |
Used for the created-by label (default; see owner_email_env_vars) |
GC_BATCH_* |
Override any setting; see Settings file |
Container Environment Variables
When you create a job with input/output mounts, gc-batch automatically sets these environment variables in your container:
| Variable | Description | Example |
|---|---|---|
INPUT_DIR |
Path to input mount | /mnt/input |
OUTPUT_DIR |
Path to output mount | /mnt/output |
Use these in your application:
import os
input_dir = os.environ.get("INPUT_DIR", "/mnt/input")
output_dir = os.environ.get("OUTPUT_DIR", "/mnt/output")
# Read from input
with open(f"{input_dir}/data.csv") as f:
data = f.read()
# Write to output
with open(f"{output_dir}/results.json", "w") as f:
f.write(results)
Override Defaults
Change Location and Project
# Use different location and project
gc-batch --location us-west1 --project-id my-other-project list-jobs
# Create job in different region
gc-batch --location europe-west1 create \
--job-name eu-job \
--docker-image python:3.12 \
--command "python main.py"
Change Machine Type
# Use a larger machine
gc-batch create \
--job-name big-job \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py" \
--machine-type n2-standard-8
# Use a high-memory machine
gc-batch create \
--job-name memory-intensive \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py" \
--machine-type n2-highmem-4
Change Disk Settings
# Larger boot disk
gc-batch create \
--job-name big-disk \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py" \
--boot-disk-size 100
# Use SSD boot disk
gc-batch create \
--job-name fast-boot \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py" \
--boot-disk-type pd-ssd
Provisioning Models
Control cost vs availability with provisioning models:
| Model | Cost | Availability | Use Case |
|---|---|---|---|
STANDARD |
Highest | Guaranteed | Production, time-sensitive jobs |
SPOT |
60-91% cheaper | May be preempted | Fault-tolerant batch jobs |
PREEMPTIBLE |
Similar to SPOT | May be preempted | Legacy (use SPOT instead) |
Using SPOT for Cost Savings
gc-batch create \
--job-name cost-optimized \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/process.py" \
--provisioning-model SPOT
When to use SPOT
Use SPOT VMs for:
- Batch processing that can be retried
- Development and testing
- Non-urgent data processing
- Any workload that can handle interruptions
SPOT Limitations
SPOT VMs can be preempted (terminated) at any time if Google Cloud needs the capacity. Your job should be able to handle restarts or save progress periodically.
Custom Environment Variables
Pass custom environment variables to your container:
gc-batch create \
--job-name env-job \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py" \
--env "DATABASE_URL=postgres://...,API_KEY=abc123,DEBUG=true"
Access in your container:
import os
database_url = os.environ.get("DATABASE_URL")
api_key = os.environ.get("API_KEY")
debug = os.environ.get("DEBUG", "false").lower() == "true"
Settings file
Values that differ between deployments are resolved in this order, highest precedence first:
- explicit arguments (CLI options, or constructor arguments in the Python API)
GC_BATCH_*environment variables- a TOML config file
- neutral defaults
The config file is discovered from --config-file, then $GC_BATCH_CONFIG_FILE,
then ./gc-batch.toml, then $XDG_CONFIG_HOME/gc-batch/config.toml.
# gc-batch.toml
default_project_id = "my-project"
job_name_prefix = "" # prepended to every job name
created_using_label = "gc-batch" # value of the created-using label
# Environment variables checked, in order, for the created-by label
owner_email_env_vars = ["USER"]
# Additional named job profiles, usable via --job-profile
[job_profiles.my-vpc]
network = "global/networks/my-network"
subnetwork = "regions/us-central1/subnetworks/my-subnet"
use_private_address = true
regions = ["us-central1"]
Run the config-display command to see the resolved settings and debug precedence.
Job profiles
A job profile is a named bundle of network and placement settings applied with
--job-profile <name>. Passing an unknown name is an error that lists the
available profiles.
The built-in all-of-us profile
gc-batch ships one built-in profile for the
All of Us Researcher Workbench, whose jobs run inside
a restricted VPC. --job-profile all-of-us applies:
network=global/networks/networksubnetwork=regions/us-central1/subnetworks/subnetworkuse_private_address=trueregions=["us-central1"]service_accounttaken from your currentgcloudauthenticationcloud_logging_unreadable=true
gc-batch --project-id my-project create \
--job-profile all-of-us \
--job-name my-job \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py"
Set --logs-bucket in the All of Us workbench
Workspace service accounts cannot read Cloud Logging, so a job that writes its
logs there produces logs nobody in the workspace can read. Because
cloud_logging_unreadable is true for this profile, create warns when
--logs-bucket is missing. Heed it: the log destination is fixed when the job is
created, so the only remedy afterwards is to recreate the job.
Declaring that Cloud Logging is unreadable
Set cloud_logging_unreadable = true on any profile whose environment cannot read
Cloud Logging, and create will warn when --logs-bucket is omitted:
[job_profiles.locked-down]
network = "global/networks/my-network"
use_private_address = true
cloud_logging_unreadable = true
The warning is advisory — it goes to stderr and does not stop the job from being created.