Example Workflows
Real-world examples of using gc-batch for common scenarios.
The examples below use the CLI; the same workflows from Python at the end of this page show the library equivalents. See the Python API guide for the full library surface.
Data Science Workflow
A typical workflow for running a data analysis job:
1. Create the Job
gc-batch create \
--job-name data-analysis-001 \
--docker-image gcr.io/my-project/data-science:latest \
--command "python /app/analyze.py" \
--args "--input /mnt/input/data.csv --output /mnt/output/results/" \
--input-bucket my-bucket/datasets \
--output-bucket my-bucket/results \
--labels "team=data-science,project=user-analysis,priority=high" \
--machine-type n2-standard-8 \
--boot-disk-size 100
2. Monitor the Job
# Check status
gc-batch status --job-name data-analysis-001-1234567890
# Watch for completion (run periodically)
gc-batch list-my-jobs --status RUNNING
3. Check Logs
# View logs in console
gc-batch logs print --job-name data-analysis-001-1234567890
# Or open in Cloud Console for better viewing
gc-batch logs url --job-name data-analysis-001-1234567890
4. Retrieve Results
Your results will be in the output bucket: gs://my-bucket/results/
Production Monitoring
Commands for monitoring production batch jobs:
List All Production Jobs
gc-batch list-jobs --labels "environment=production"
Find Running Production Jobs
gc-batch list-jobs --labels "environment=production" --status RUNNING
Find Failed Production Jobs
# Failed in the last day
gc-batch list-jobs --labels "environment=production" --status FAILED --since 1d
# Failed in the last week
gc-batch list-jobs --labels "environment=production" --status FAILED --since 7d
Debug a Failed Job
# Get status with error details
gc-batch status --job-name production-job-1234567890 --full
# View error logs
gc-batch logs url --job-name production-job-1234567890 --severity ERROR
Team Management
Commands for managing jobs across a team:
List All Team Jobs
gc-batch list-jobs --labels "team=bioinformatics"
List Your Own Jobs
gc-batch list-my-jobs --since 1d
List Failed Jobs for Your Team
gc-batch list-jobs --labels "team=bioinformatics" --status FAILED --since 7d
List Running Jobs for Your Team
gc-batch list-jobs --labels "team=bioinformatics" --status RUNNING
High-Performance Processing
For jobs requiring fast I/O:
Using Local SSD
gc-batch create \
--job-name high-io-processing \
--docker-image gcr.io/my-project/processor:latest \
--command "python /app/process.py" \
--machine-type n2-standard-8 \
--local-ssd-size-gb 375 \
--local-ssd-mount-path /mnt/fast \
--input-bucket my-bucket/large-dataset \
--output-bucket my-bucket/processed \
--labels "use-case=high-io,team=data-engineering"
Using LSSD Machine Types
gc-batch create \
--job-name lssd-processing \
--docker-image gcr.io/my-project/processor:latest \
--command "python /app/process.py" \
--machine-type c4-standard-8-lssd \
--local-ssd-mount-path /mnt/fast \
--labels "use-case=high-io,machine=lssd"
Cost-Optimized Batch Processing
For non-urgent jobs where cost is a priority:
Using SPOT VMs
gc-batch create \
--job-name nightly-batch \
--docker-image gcr.io/my-project/batch-processor:latest \
--command "python /app/batch_process.py" \
--provisioning-model SPOT \
--machine-type e2-standard-4 \
--input-bucket my-bucket/daily-data \
--output-bucket my-bucket/daily-results \
--labels "cost-optimization=spot,schedule=nightly"
ML Training Pipeline
Example for machine learning training:
Training Job
gc-batch create \
--job-name ml-training-v1 \
--docker-image gcr.io/my-project/ml-trainer:latest \
--command "python /app/train.py" \
--args "--epochs 100 --batch-size 32 --model-name resnet50" \
--input-bucket my-ml-bucket/training-data \
--output-bucket my-ml-bucket/models/v1 \
--machine-type n2-standard-16 \
--boot-disk-size 200 \
--labels "team=ml,project=image-classification,version=v1"
Check Training Progress
# View training logs
gc-batch logs print --job-name ml-training-v1-1234567890
# Or in Cloud Console
gc-batch logs url --job-name ml-training-v1-1234567890
Genomics Processing
Example for bioinformatics workloads:
Genome Sequencing Job
gc-batch create \
--job-name genome-sequencing-sample001 \
--docker-image gcr.io/my-project/sequencing:latest \
--command "python /app/sequence.py" \
--args "--sample-id SAMPLE_001 --reference hg38" \
--input-bucket my-genomics-bucket/raw-samples/SAMPLE_001 \
--output-bucket my-genomics-bucket/processed/SAMPLE_001 \
--machine-type n2-highmem-8 \
--boot-disk-size 500 \
--local-ssd-size-gb 375 \
--local-ssd-mount-path /mnt/scratch \
--labels "team=bioinformatics,project=genomics,sample=SAMPLE_001"
Batch Job with Custom Environment
Passing configuration via environment variables:
gc-batch create \
--job-name configured-job \
--docker-image gcr.io/my-project/my-app:latest \
--command "python /app/main.py" \
--env "DATABASE_URL=postgres://...,LOG_LEVEL=DEBUG,MAX_WORKERS=4" \
--input-bucket my-bucket/config \
--output-bucket my-bucket/output \
--labels "environment=staging,config=custom"
The Same Workflows From Python
Anything above can be done from a script or notebook instead. All of these share one client:
from gc_batch import BatchClientConfig, BatchJobConfig, GCBatchClient, JobRequest
client = GCBatchClient(BatchClientConfig(project_id="my-project", location="us-central1"))
Data Science Workflow
Submit, wait, then act on the outcome — the part that is awkward to script around the CLI:
import time
from gc_batch.utils import is_job_finished
job = client.create_job(
JobRequest(
job_name="data-analysis-001",
docker_image="gcr.io/my-project/data-science:latest",
command="python /app/analyze.py",
args="--input /mnt/input/data.csv --output /mnt/output/results/",
config=BatchJobConfig(
machine_type="n2-standard-8",
boot_disk_type="pd-balanced",
boot_disk_size=100,
input_bucket="my-bucket/datasets",
output_bucket="my-bucket/results",
),
labels={"team": "data-science", "project": "user-analysis", "priority": "high"},
)
)
short_name = job.name.split("/")[-1]
while not is_job_finished(job := client.get_job(short_name)):
time.sleep(30)
if job.status.state.name == "SUCCEEDED":
print("Results in gs://my-bucket/results/")
else:
print(client.get_failure_message(job))
client.batch_logging.print_logs_for_job(job, severity="ERROR")
Production Monitoring
from datetime import datetime, timedelta, timezone
running = client.list_jobs(labels={"environment": "production"}, status="RUNNING")
print(f"{len(running)} production jobs running")
failures = client.list_jobs(
labels={"environment": "production"},
status="FAILED",
since_time=datetime.now(timezone.utc) - timedelta(days=1),
)
for job in failures:
print(f"\n{job.name.split('/')[-1]}")
print(client.get_failure_message(job))
print(client.get_cloud_logging_url(job, severity="ERROR"))
This is the natural hook for an alerting job: run it on a schedule and post the failure summaries and log URLs wherever your team watches.
Cost-Optimized Batch Processing
job = client.create_job(
JobRequest(
job_name="nightly-batch",
docker_image="gcr.io/my-project/batch-processor:latest",
command="python /app/batch_process.py",
config=BatchJobConfig(
machine_type="e2-standard-4",
boot_disk_type="pd-balanced",
provisioning_model="SPOT",
input_bucket="my-bucket/daily-data",
output_bucket="my-bucket/daily-results",
),
labels={"cost-optimization": "spot", "schedule": "nightly"},
)
)
ML Hyperparameter Sweep
Submitting a job per configuration is where the library pulls ahead of the CLI — the sweep is a loop, not a generated shell script:
GRID = [
{"epochs": 50, "batch_size": 32},
{"epochs": 50, "batch_size": 64},
{"epochs": 100, "batch_size": 32},
{"epochs": 100, "batch_size": 64},
]
config = BatchJobConfig(
machine_type="n2-standard-16",
boot_disk_type="pd-balanced",
boot_disk_size=200,
input_bucket="my-ml-bucket/training-data",
output_bucket="my-ml-bucket/models/sweep",
)
for index, params in enumerate(GRID):
client.create_job(
JobRequest(
job_name=f"ml-training-sweep-{index}",
docker_image="gcr.io/my-project/ml-trainer:latest",
command="python /app/train.py",
args=(
f"--epochs {params['epochs']} "
f"--batch-size {params['batch_size']} "
f"--model-name resnet50"
),
config=config,
labels={
"team": "ml",
"project": "image-classification",
"sweep": "resnet50-v1",
"epochs": str(params["epochs"]),
"batch-size": str(params["batch_size"]),
},
)
)
Then watch the whole sweep with a single call, since every job carries the same
sweep label:
sweep_jobs = client.list_jobs(labels={"sweep": "resnet50-v1"})
for job in sweep_jobs:
print(job.name.split("/")[-1], job.status.state.name)