gc_batch.models.job_request
Job request and configuration models for Google Cloud Batch.
This module contains Pydantic models for configuring and creating Batch jobs, including machine type helpers, job configuration, and job requests.
BatchBootDiskType
Bases: CustomStrEnum
Supported boot disk type strings for Batch jobs.
This is the union of disk types supported on at least one machine generation;
use :meth:MachineTypeHelper.disk_type_supported to check compatibility with
a specific machine type.
Source code in src/gc_batch/models/job_request.py
16 17 18 19 20 21 22 23 24 25 26 27 28 29 | |
HYPERDISK_BALANCED = 'hyperdisk-balanced'
class-attribute
instance-attribute
HYPERDISK_BALANCED_HIGH_AVAILABILITY = 'hyperdisk-balanced-high-availability'
class-attribute
instance-attribute
HYPERDISK_EXTREME = 'hyperdisk-extreme'
class-attribute
instance-attribute
PD_BALANCED = 'pd-balanced'
class-attribute
instance-attribute
PD_SSD = 'pd-ssd'
class-attribute
instance-attribute
PD_STANDARD = 'pd-standard'
class-attribute
instance-attribute
BatchJobConfig
Bases: BaseModel
Configuration for a Google Cloud Batch job.
This model contains all the configuration options for creating a Batch job, including machine specifications, storage, networking, and environment settings.
Attributes:
| Name | Type | Description |
|---|---|---|
default_task_count |
int
|
Number of tasks in the job (default: 1). |
default_parallelism |
int
|
Maximum tasks to run in parallel (default: 1). |
machine_type |
str
|
GCP machine type (e.g., "e2-standard-2", "n2-standard-4"). |
boot_disk_size |
int
|
Boot disk size in GB (default: 30). |
boot_disk_type |
BatchBootDiskType
|
Boot disk type (:class: |
input_bucket |
str | None
|
GCS bucket path for input data (without gs:// prefix). |
input_dir |
str | None
|
Mount point for input data in container (default: "/mnt/input"). |
input_billing_project |
str | None
|
Billing project for requester-pays input buckets. |
output_bucket |
str | None
|
GCS bucket path for output data (without gs:// prefix). |
output_dir |
str | None
|
Mount point for output data in container (default: "/mnt/output"). |
logs_bucket |
str | None
|
GCS bucket path (without gs:// prefix) to write job logs to. When set, logs go to this bucket path instead of Cloud Logging. Include a subfolder in the path to separate logs per job. |
logs_billing_project |
str | None
|
Billing project for requester-pays logs buckets. |
local_ssd_size_gb |
int | None
|
Size of local SSD in GB (must be multiple of 375). |
local_ssd_device_name |
str | None
|
Device name for local SSD (default: "local-ssd-0"). |
local_ssd_mount_path |
str | None
|
Mount path for local SSD (default: "/mnt/local_ssd"). |
provisioning_model |
BatchProvisioningModel | None
|
VM provisioning model (STANDARD, SPOT, or PREEMPTIBLE). |
network |
str | None
|
VPC network path for the VM. |
subnetwork |
str | None
|
Subnetwork path for the VM. |
service_account |
str | None
|
Service account email to use for the job. |
use_private_address |
bool
|
Whether to use private IP (no external IP). |
regions |
list[str] | None
|
List of allowed regions for job placement. |
zones |
list[str] | None
|
List of allowed zones for job placement. |
user_env_dict |
dict[str, str] | None
|
Custom environment variables for the container. |
Example
from gc_batch import BatchJobConfig
# Basic configuration
config = BatchJobConfig(
machine_type="n2-standard-4",
boot_disk_type="pd-balanced",
boot_disk_size=50,
)
# Configuration with GCS mounts
config = BatchJobConfig(
machine_type="n2-standard-8",
boot_disk_type="pd-ssd",
input_bucket="my-bucket/input-data",
output_bucket="my-bucket/output-data",
)
# Configuration with local SSD
config = BatchJobConfig(
machine_type="n2-standard-4",
boot_disk_type="pd-balanced",
local_ssd_size_gb=375,
local_ssd_mount_path="/mnt/fast",
)
Source code in src/gc_batch/models/job_request.py
179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 | |
boot_disk_size = Field(default=30)
class-attribute
instance-attribute
boot_disk_type
instance-attribute
cores = Field(default=1)
class-attribute
instance-attribute
default_parallelism = Field(default=1)
class-attribute
instance-attribute
default_task_count = Field(default=1)
class-attribute
instance-attribute
input_billing_project = Field(default=None, description='Billing project for input bucket access')
class-attribute
instance-attribute
input_bucket = Field(default=None)
class-attribute
instance-attribute
input_dir = Field(default=(Constants.INPUT_DIR))
class-attribute
instance-attribute
local_ssd_device_name = Field(default='local-ssd-0', description="Device name for the local SSD (default: 'local-ssd-0')")
class-attribute
instance-attribute
local_ssd_mount_path = Field(default=None, description="Mount path for the local SSD in the container (default: '/mnt/local_ssd')")
class-attribute
instance-attribute
local_ssd_size_gb = Field(default=None, description='Size of local SSD in GB (must be multiple of 375 GB). If specified, a local SSD will be attached.')
class-attribute
instance-attribute
logs_billing_project = Field(default=None, description='Billing project for requester-pays logs bucket access')
class-attribute
instance-attribute
logs_bucket = Field(default=None, description="GCS bucket path (without the gs:// prefix) where job logs will be written, e.g. 'my-bucket/batch-logs'. When set, logs are written to this bucket path instead of Cloud Logging; include a subfolder to separate logs per job. Useful when Cloud Logging is not accessible. Note: Batch supports a single log destination, so enabling this disables Cloud Logging for the job.")
class-attribute
instance-attribute
machine_type
instance-attribute
network = Field(default=None, description="Network to use (e.g., 'global/networks/network')")
class-attribute
instance-attribute
output_bucket = Field(default=None)
class-attribute
instance-attribute
output_dir = Field(default=(Constants.OUTPUT_DIR))
class-attribute
instance-attribute
output_location = Field(default=None)
class-attribute
instance-attribute
provisioning_model = Field(default=None, description="Provisioning model: 'STANDARD', 'SPOT', or 'PREEMPTIBLE'. SPOT is recommended for cost savings. Accepts string input which is converted to enum.")
class-attribute
instance-attribute
ram_gb = Field(default=4)
class-attribute
instance-attribute
regions = Field(default=None, description='List of regions to use')
class-attribute
instance-attribute
service_account = Field(default=None, description='Service account email to use')
class-attribute
instance-attribute
subnetwork = Field(default=None, description="Subnetwork to use (e.g., 'regions/us-central1/subnetworks/subnetwork')")
class-attribute
instance-attribute
use_private_address = Field(default=False, description='Use private IP address (no external IP)')
class-attribute
instance-attribute
user_env_dict = Field(default=None)
class-attribute
instance-attribute
zones = Field(default=None, description='List of zones to use')
class-attribute
instance-attribute
apply_profile(profile)
Apply a job profile's networking and VM settings to this config.
This is what gc-batch create --job-profile does, exposed for library
callers. Fields the profile leaves unset are left untouched, so a profile
can be applied over a config that already has other settings.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
profile
|
JobProfile
|
The :class: |
required |
Returns:
| Type | Description |
|---|---|
BatchJobConfig
|
This config, mutated in place, to allow chaining after construction. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the profile sets |
Example
from gc_batch import BatchJobConfig, GCBatchSettings
settings = GCBatchSettings.load()
config = BatchJobConfig(
machine_type="n2-standard-4",
boot_disk_type="pd-balanced",
).apply_profile(settings.job_profiles["all-of-us"])
Source code in src/gc_batch/models/job_request.py
329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 | |
validate_boot_disk_type(v)
classmethod
Convert string input to BatchBootDiskType enum before validation.
Source code in src/gc_batch/models/job_request.py
321 322 323 324 325 326 327 | |
validate_provisioning_model(v)
classmethod
Convert string input to BatchProvisioningModel enum before validation.
Source code in src/gc_batch/models/job_request.py
311 312 313 314 315 316 317 318 319 | |
JobRequest
Bases: BaseModel
Request to create a new Google Cloud Batch job.
This model represents a complete job creation request, combining the job metadata with the job configuration.
Attributes:
| Name | Type | Description |
|---|---|---|
job_name |
str
|
Name for the job (will be timestamped, and prefixed if
|
docker_image |
str
|
Docker image URI to use for the job container. |
command |
str
|
Command to run inside the container. |
args |
str
|
Additional arguments to pass to the command (optional). |
config |
BatchJobConfig
|
BatchJobConfig with machine and storage settings. |
labels |
dict[str, str] | None
|
Custom labels to attach to the job for filtering and organization. |
Example
from gc_batch import JobRequest, BatchJobConfig
config = BatchJobConfig(
machine_type="n2-standard-4",
boot_disk_type="pd-balanced",
input_bucket="my-bucket/data",
output_bucket="my-bucket/results",
)
request = JobRequest(
job_name="data-analysis",
docker_image="gcr.io/my-project/analyzer:latest",
command="python /app/analyze.py",
args="--input /mnt/input --output /mnt/output",
config=config,
labels={"team": "data-science", "environment": "production"},
)
Source code in src/gc_batch/models/job_request.py
378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 | |
args = Field(default='')
class-attribute
instance-attribute
command = Field(default="echo 'Hello, World!'")
class-attribute
instance-attribute
config
instance-attribute
docker_image = Field(default='ubuntu:latest')
class-attribute
instance-attribute
job_name = Field(default='default')
class-attribute
instance-attribute
labels = Field(default=None, description='Custom labels for the Batch job')
class-attribute
instance-attribute
MachineTypeHelper
Helper class for working with GCP machine types.
Provides utilities for determining compatible disk types, checking machine type generations, and identifying LSSD (Local SSD) machine types.
Example
from gc_batch import MachineTypeHelper
# Check if a machine type is LSSD
is_lssd = MachineTypeHelper.is_lssd_machine_type("c4-standard-8-lssd")
# Returns: True
# Get supported disk types for a machine type
disk_types = MachineTypeHelper.supported_disk_types("n2-standard-4")
# Returns: ["hyperdisk-balanced", "hyperdisk-balanced-high-availability", ...]
# Get the default disk type
default = MachineTypeHelper.get_default_disk_type("e2-standard-2")
# Returns: "pd-standard"
Source code in src/gc_batch/models/job_request.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 | |
disk_type_supported(disk_type, type_name)
staticmethod
Check if a disk type is supported for a given machine type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
disk_type
|
str
|
The disk type to check (e.g., "pd-balanced"). |
required |
type_name
|
str
|
The machine type name (e.g., "n2-standard-4"). |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if the disk type is supported, False otherwise. |
Source code in src/gc_batch/models/job_request.py
93 94 95 96 97 98 99 100 101 102 103 104 | |
get_default_disk_type(type_name)
staticmethod
Get the default disk type for a machine type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
type_name
|
str
|
The machine type name (e.g., "n2-standard-4"). |
required |
Returns:
| Type | Description |
|---|---|
str
|
The default (first) supported disk type for this machine type. |
Source code in src/gc_batch/models/job_request.py
121 122 123 124 125 126 127 128 129 130 131 | |
is_lssd_machine_type(type_name)
staticmethod
Check if a machine type has pre-attached Local SSDs.
LSSD machine types (e.g., "c4-standard-8-lssd") come with Local SSDs automatically attached and cannot have additional SSDs manually attached.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
type_name
|
str
|
The machine type name to check. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if the machine type is an LSSD type, False otherwise. |
Source code in src/gc_batch/models/job_request.py
106 107 108 109 110 111 112 113 114 115 116 117 118 119 | |
is_mid_generation_machine_type(type_name)
staticmethod
Check if a machine type is generation 3 or newer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
type_name
|
str
|
The machine type name (e.g., "n3-standard-4"). |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if generation 3+, False otherwise. |
Source code in src/gc_batch/models/job_request.py
145 146 147 148 149 150 151 152 153 154 155 | |
is_new_generation_machine_type(type_name)
staticmethod
Check if a machine type is generation 4 or newer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
type_name
|
str
|
The machine type name (e.g., "c4-standard-8"). |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if generation 4+, False otherwise. |
Source code in src/gc_batch/models/job_request.py
133 134 135 136 137 138 139 140 141 142 143 | |
supported_disk_types(type_name)
staticmethod
Get the list of supported disk types for a machine type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
type_name
|
str
|
The GCP machine type name (e.g., "n2-standard-4"). |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
List of supported disk type strings for this machine type. |
Note
- Generation 4+ machines (c4, m4, etc.) only support hyperdisk types.
- Generation 3 machines support both hyperdisk and pd-* types.
- Older generations only support pd-* types.
Source code in src/gc_batch/models/job_request.py
56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 | |