I/O Utilities
I/O utilities for GRIB2, NetCDF, and SCHISM grid data.
- class nos_utils.io.GRIBExtractor[source]
Bases:
ABCAbstract GRIB2 extraction interface.
- abstractmethod extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float]) numpy.ndarray | None[source]
Extract a single variable from a GRIB2 file, subsetted to domain.
- Parameters:
grib_file – Path to GRIB2 file
variable – GRIB2 variable name (e.g., “UGRD”)
level – Level string (e.g., “10 m above ground”)
domain – (lon_min, lon_max, lat_min, lat_max)
- Returns:
2D numpy array (ny, nx) or None if extraction failed
- abstractmethod get_grid(grib_file: Path, domain: Tuple[float, float, float, float]) Tuple[numpy.ndarray, numpy.ndarray][source]
Get lon/lat coordinate arrays for the subsetted domain.
- Returns:
(lons_1d, lats_1d) arrays
- regrid_to_latlon(grib_file: Path, domain: Tuple[float, float, float, float], dx: float, output_path: Path) Path | None[source]
Regrid a GRIB2 file from native projection to regular lat/lon.
Critical for HRRR (Lambert Conformal) — SCHISM cannot handle LCC grids.
- Parameters:
grib_file – Input GRIB2 file (any projection)
domain – Target domain bounds
dx – Target resolution in degrees
output_path – Path for regridded output
- Returns:
Path to regridded GRIB2 file, or None if not supported
- class nos_utils.io.Wgrib2Extractor(wgrib2_path: str = 'wgrib2')[source]
Bases:
GRIBExtractorProduction GRIB2 extractor using wgrib2 subprocess.
wgrib2 handles all grid projections, variable matching, and domain subsetting natively. This is the preferred backend for operational use.
Auto-detects wgrib2 with IPOLATES support (needed for -new_grid regridding). Prefers spack-stack wgrib2 over system wgrib2 when available.
- WGRIB2_SEARCH_PATHS = ['/opt/spack-stack/spack-stack-1.9.2/envs/ufs-wm-env/install/gcc/13.3.1/wgrib2-3.6.0-yu365ku/bin/wgrib2', '/opt/spack-stack/*/envs/*/install/*/wgrib2-*/bin/wgrib2']
- extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float], skip_subset: bool = False) numpy.ndarray | None[source]
Extract a variable from GRIB2 file.
- Parameters:
skip_subset – If True, skip -small_grib domain subsetting. Use when the file is already regridded to the target domain (avoids off-by-one trimming from floating-point boundary matching).
- get_grid(grib_file: Path, domain: Tuple[float, float, float, float]) Tuple[numpy.ndarray, numpy.ndarray][source]
Get lon/lat coordinate arrays for the subsetted domain.
- Returns:
(lons_1d, lats_1d) arrays
- regrid_to_latlon(grib_file: Path, domain: Tuple[float, float, float, float], dx: float, output_path: Path, match_pattern: str | None = None) Path | None[source]
Regrid GRIB2 from native projection (e.g., Lambert Conformal) to regular lat/lon.
Uses: wgrib2 -new_grid_winds earth -new_grid latlon lon0:nx:dx lat0:ny:dy The -new_grid_winds earth flag correctly rotates grid-relative winds.
For wind rotation to work, U and V must be processed together in the same command (wgrib2 pairs them automatically when both are present).
- Parameters:
grib_file – Input GRIB2 file
domain – (lon_min, lon_max, lat_min, lat_max)
dx – Target resolution in degrees
output_path – Path for regridded output
match_pattern – Optional wgrib2 -match regex to select variables. For HRRR, use a pattern that includes BOTH UGRD and VGRD: “:(UGRD:10 m|VGRD:10 m|TMP:2 m|SPFH:2 m|MSLMA|PRATE|DSWRF|DLWRF):”
- class nos_utils.io.CfgribExtractor[source]
Bases:
GRIBExtractorDevelopment GRIB2 extractor using cfgrib/xarray.
No external wgrib2 binary needed. Useful for local development and testing. Slower than wgrib2 and does not support all projection types.
- extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float]) numpy.ndarray | None[source]
Extract a single variable from a GRIB2 file, subsetted to domain.
- Parameters:
grib_file – Path to GRIB2 file
variable – GRIB2 variable name (e.g., “UGRD”)
level – Level string (e.g., “10 m above ground”)
domain – (lon_min, lon_max, lat_min, lat_max)
- Returns:
2D numpy array (ny, nx) or None if extraction failed
- nos_utils.io.get_extractor(wgrib2_path: str = 'wgrib2') GRIBExtractor[source]
Auto-detect and return the best available GRIB2 extractor.
Prefers wgrib2 (faster, handles all projections) over cfgrib.
- Raises:
RuntimeError – If neither wgrib2 nor cfgrib is available
- class nos_utils.io.SchismGrid(filepath: Path, n_nodes: int, n_elements: int, node_lons: numpy.ndarray, node_lats: numpy.ndarray, node_depths: numpy.ndarray, open_boundaries: List[OpenBoundary] = <factory>, n_open_boundary_nodes: int = 0)[source]
Bases:
objectParsed SCHISM grid with boundary information.
- Usage:
grid = SchismGrid.read(“secofs.hgrid.ll”) bnd_nodes = grid.open_boundary_nodes() # all open boundary node coords
- node_lons: numpy.ndarray
- node_lats: numpy.ndarray
- node_depths: numpy.ndarray
- open_boundaries: List[OpenBoundary]
- classmethod read(filepath, read_boundaries: bool = True) SchismGrid[source]
Read SCHISM grid file (hgrid.gr3 or hgrid.ll).
- Parameters:
filepath – Path to grid file
read_boundaries – If True, parse open boundary sections
- Returns:
SchismGrid instance
- open_boundary_nodes() Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, List[int]][source]
Get all open boundary node coordinates concatenated.
- Returns:
(lons, lats, depths, node_ids) — all open boundary nodes across all segments
- obc_nodes_from_ctl(ctl_path: Path) Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, List[int]][source]
Read boundary node IDs from obc.ctl file (exact match with Fortran).
The obc.ctl has the precise 1,488 nodes used by gen_3Dth_from_hycom, which may differ slightly from hgrid.ll boundary section.
- Parameters:
ctl_path – Path to {ofs}.obc.ctl
- Returns:
(lons, lats, depths, node_ids) — from obc.ctl SECTION 2
- static read_gr3_values(filepath) Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, numpy.ndarray][source]
Read node values from a gr3-format file (e.g., nudge weight files).
This is a lightweight reader that only parses the node section, skipping elements and boundaries. Suitable for large gr3 files (300+ MB) where only the per-node values are needed.
The gr3 format has lines:
node_id lon lat value [extras...]after the 2-line header.- Parameters:
filepath – Path to gr3 file
- Returns:
(node_ids_1based, lons, lats, values) — all as numpy arrays
- class nos_utils.io.OpenBoundary(index: int, node_ids: List[int], lons: numpy.ndarray, lats: numpy.ndarray, depths: numpy.ndarray)[source]
Bases:
objectA single open boundary segment.
- lons: numpy.ndarray
- lats: numpy.ndarray
- depths: numpy.ndarray
GRIB2 extraction
GRIB2 data extraction backends.
Two backends: - Wgrib2Extractor: Production — uses wgrib2 subprocess (fast, handles all grids) - CfgribExtractor: Development — uses cfgrib/xarray (no external binary needed)
Usage:
extractor = get_extractor() # auto-detect best available
data, lons, lats = extractor.extract("gfs.grib2", "UGRD", "10 m above ground",
domain=(-88, -63, 17, 40))
- class nos_utils.io.grib_extract.GRIBExtractor[source]
Bases:
ABCAbstract GRIB2 extraction interface.
- abstractmethod extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float]) numpy.ndarray | None[source]
Extract a single variable from a GRIB2 file, subsetted to domain.
- Parameters:
grib_file – Path to GRIB2 file
variable – GRIB2 variable name (e.g., “UGRD”)
level – Level string (e.g., “10 m above ground”)
domain – (lon_min, lon_max, lat_min, lat_max)
- Returns:
2D numpy array (ny, nx) or None if extraction failed
- abstractmethod get_grid(grib_file: Path, domain: Tuple[float, float, float, float]) Tuple[numpy.ndarray, numpy.ndarray][source]
Get lon/lat coordinate arrays for the subsetted domain.
- Returns:
(lons_1d, lats_1d) arrays
- regrid_to_latlon(grib_file: Path, domain: Tuple[float, float, float, float], dx: float, output_path: Path) Path | None[source]
Regrid a GRIB2 file from native projection to regular lat/lon.
Critical for HRRR (Lambert Conformal) — SCHISM cannot handle LCC grids.
- Parameters:
grib_file – Input GRIB2 file (any projection)
domain – Target domain bounds
dx – Target resolution in degrees
output_path – Path for regridded output
- Returns:
Path to regridded GRIB2 file, or None if not supported
- class nos_utils.io.grib_extract.Wgrib2Extractor(wgrib2_path: str = 'wgrib2')[source]
Bases:
GRIBExtractorProduction GRIB2 extractor using wgrib2 subprocess.
wgrib2 handles all grid projections, variable matching, and domain subsetting natively. This is the preferred backend for operational use.
Auto-detects wgrib2 with IPOLATES support (needed for -new_grid regridding). Prefers spack-stack wgrib2 over system wgrib2 when available.
- WGRIB2_SEARCH_PATHS = ['/opt/spack-stack/spack-stack-1.9.2/envs/ufs-wm-env/install/gcc/13.3.1/wgrib2-3.6.0-yu365ku/bin/wgrib2', '/opt/spack-stack/*/envs/*/install/*/wgrib2-*/bin/wgrib2']
- extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float], skip_subset: bool = False) numpy.ndarray | None[source]
Extract a variable from GRIB2 file.
- Parameters:
skip_subset – If True, skip -small_grib domain subsetting. Use when the file is already regridded to the target domain (avoids off-by-one trimming from floating-point boundary matching).
- get_grid(grib_file: Path, domain: Tuple[float, float, float, float]) Tuple[numpy.ndarray, numpy.ndarray][source]
Get lon/lat coordinate arrays for the subsetted domain.
- Returns:
(lons_1d, lats_1d) arrays
- regrid_to_latlon(grib_file: Path, domain: Tuple[float, float, float, float], dx: float, output_path: Path, match_pattern: str | None = None) Path | None[source]
Regrid GRIB2 from native projection (e.g., Lambert Conformal) to regular lat/lon.
Uses: wgrib2 -new_grid_winds earth -new_grid latlon lon0:nx:dx lat0:ny:dy The -new_grid_winds earth flag correctly rotates grid-relative winds.
For wind rotation to work, U and V must be processed together in the same command (wgrib2 pairs them automatically when both are present).
- Parameters:
grib_file – Input GRIB2 file
domain – (lon_min, lon_max, lat_min, lat_max)
dx – Target resolution in degrees
output_path – Path for regridded output
match_pattern – Optional wgrib2 -match regex to select variables. For HRRR, use a pattern that includes BOTH UGRD and VGRD: “:(UGRD:10 m|VGRD:10 m|TMP:2 m|SPFH:2 m|MSLMA|PRATE|DSWRF|DLWRF):”
- class nos_utils.io.grib_extract.CfgribExtractor[source]
Bases:
GRIBExtractorDevelopment GRIB2 extractor using cfgrib/xarray.
No external wgrib2 binary needed. Useful for local development and testing. Slower than wgrib2 and does not support all projection types.
- extract(grib_file: Path, variable: str, level: str, domain: Tuple[float, float, float, float]) numpy.ndarray | None[source]
Extract a single variable from a GRIB2 file, subsetted to domain.
- Parameters:
grib_file – Path to GRIB2 file
variable – GRIB2 variable name (e.g., “UGRD”)
level – Level string (e.g., “10 m above ground”)
domain – (lon_min, lon_max, lat_min, lat_max)
- Returns:
2D numpy array (ny, nx) or None if extraction failed
- nos_utils.io.grib_extract.get_extractor(wgrib2_path: str = 'wgrib2') GRIBExtractor[source]
Auto-detect and return the best available GRIB2 extractor.
Prefers wgrib2 (faster, handles all projections) over cfgrib.
- Raises:
RuntimeError – If neither wgrib2 nor cfgrib is available
NetCDF utilities
Shared NetCDF utilities for nos-utils.
Common helpers for time axis creation, fill value handling, variable subsetting, and dimension naming conventions.
- nos_utils.io.netcdf_utils.read_time_axis(filepath: Path) Tuple[List[float], str][source]
Read time values and units from a NetCDF file.
- Returns:
(time_values, units_string)
- nos_utils.io.netcdf_utils.validate_monotonic(time_values: List[float], label: str = 'time') bool[source]
Check that a time axis is strictly monotonically increasing.
Raises ValueError if not monotonic.
- nos_utils.io.netcdf_utils.replace_fill_values(data: numpy.ndarray, threshold: float = 10000.0, fill_value: float = -9999.0) numpy.ndarray[source]
Replace extreme values (abs > threshold) with fill_value.
RTOFS data sometimes has values > 10000 for missing data.
- nos_utils.io.netcdf_utils.get_grid_dims(filepath: Path) Tuple[int, int][source]
Read (nx, ny) from a NetCDF file’s coordinate dimensions.
Checks common dimension names: longitude/latitude, lon/lat, nx_grid/ny_grid, x/y.
- nos_utils.io.netcdf_utils.subset_domain(data: numpy.ndarray, lons: numpy.ndarray, lats: numpy.ndarray, domain: Tuple[float, float, float, float]) Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray][source]
Subset a 2D field to a lon/lat bounding box.
- Parameters:
data – 2D array (ny, nx)
lons – 1D longitude array
lats – 1D latitude array
domain – (lon_min, lon_max, lat_min, lat_max)
- Returns:
(subset_data, subset_lons, subset_lats)
SCHISM grid reader
SCHISM grid reader (hgrid.gr3 / hgrid.ll).
- Parses the SCHISM unstructured grid file to extract:
Node coordinates (lon, lat, depth)
Open boundary node indices and coordinates
Land boundary information
Works with any SCHISM-based OFS (SECOFS, STOFS-3D-ATL, CREOFS, etc.)
- hgrid format:
Line 1: comment Line 2: n_elements n_nodes Lines 3 to n_nodes+2: node_id lon lat depth Lines n_nodes+3 to n_nodes+n_elements+2: element connectivity After elements: open boundary section, then land boundary section
- class nos_utils.io.schism_grid.OpenBoundary(index: int, node_ids: List[int], lons: numpy.ndarray, lats: numpy.ndarray, depths: numpy.ndarray)[source]
Bases:
objectA single open boundary segment.
- lons: numpy.ndarray
- lats: numpy.ndarray
- depths: numpy.ndarray
- class nos_utils.io.schism_grid.SchismGrid(filepath: Path, n_nodes: int, n_elements: int, node_lons: numpy.ndarray, node_lats: numpy.ndarray, node_depths: numpy.ndarray, open_boundaries: List[OpenBoundary] = <factory>, n_open_boundary_nodes: int = 0)[source]
Bases:
objectParsed SCHISM grid with boundary information.
- Usage:
grid = SchismGrid.read(“secofs.hgrid.ll”) bnd_nodes = grid.open_boundary_nodes() # all open boundary node coords
- node_lons: numpy.ndarray
- node_lats: numpy.ndarray
- node_depths: numpy.ndarray
- open_boundaries: List[OpenBoundary]
- classmethod read(filepath, read_boundaries: bool = True) SchismGrid[source]
Read SCHISM grid file (hgrid.gr3 or hgrid.ll).
- Parameters:
filepath – Path to grid file
read_boundaries – If True, parse open boundary sections
- Returns:
SchismGrid instance
- open_boundary_nodes() Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, List[int]][source]
Get all open boundary node coordinates concatenated.
- Returns:
(lons, lats, depths, node_ids) — all open boundary nodes across all segments
- obc_nodes_from_ctl(ctl_path: Path) Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, List[int]][source]
Read boundary node IDs from obc.ctl file (exact match with Fortran).
The obc.ctl has the precise 1,488 nodes used by gen_3Dth_from_hycom, which may differ slightly from hgrid.ll boundary section.
- Parameters:
ctl_path – Path to {ofs}.obc.ctl
- Returns:
(lons, lats, depths, node_ids) — from obc.ctl SECTION 2
- static read_gr3_values(filepath) Tuple[numpy.ndarray, numpy.ndarray, numpy.ndarray, numpy.ndarray][source]
Read node values from a gr3-format file (e.g., nudge weight files).
This is a lightweight reader that only parses the node section, skipping elements and boundaries. Suitable for large gr3 files (300+ MB) where only the per-node values are needed.
The gr3 format has lines:
node_id lon lat value [extras...]after the 2-line header.- Parameters:
filepath – Path to gr3 file
- Returns:
(node_ids_1based, lons, lats, values) — all as numpy arrays
SCHISM vertical grid reader
SCHISM vertical grid reader (vgrid.in).
Parses the simple vgrid.in format (from FIXofs work directory) to extract Z-levels and S-levels (sigma coordinates) for vertical interpolation.
- Format:
Line 1: nvrt kz h_s (total levels, Z-level count, S-Z transition depth) Line 2: “Z levels” Lines 3 to kz+2: level_index depth_m Line kz+3: “S levels” header Remaining: level_index sigma_value
- class nos_utils.io.schism_vgrid.SchismVgrid(nvrt: int, kz: int, h_s: float, z_levels: numpy.ndarray, sigma_levels: numpy.ndarray, node_sigma: numpy.ndarray | None = None, node_kbp: numpy.ndarray | None = None, _filepath: str | None = None)[source]
Bases:
objectParsed SCHISM vertical grid.
- z_levels: numpy.ndarray
- sigma_levels: numpy.ndarray
- node_sigma: numpy.ndarray | None = None
- node_kbp: numpy.ndarray | None = None
- load_boundary_sigma(boundary_node_ids: List[int]) None[source]
Load per-node sigma values for boundary nodes from LSC2 vgrid.in.
This reads the full file but only extracts columns for boundary nodes. Takes ~60-80 seconds for 1.68M nodes × 63 levels.
- Parameters:
boundary_node_ids – List of 1-based node IDs
- get_node_depths(node_idx: int, bottom_depth: float) numpy.ndarray[source]
Get actual depths for a specific node using per-node sigma values.
For LSC2 format where each node has its own sigma distribution. Returns depths in meters (negative = below surface).
- Parameters:
node_idx – Index into the boundary node sigma array (0-based)
bottom_depth – Positive bottom depth in meters
- Returns:
Array of depth values for valid levels (sigma > -8)
- get_depths(bottom_depth: float) numpy.ndarray[source]
Compute actual depths for a node with given bottom depth.
For deep nodes (depth > h_s): Z-levels + S-levels mapped to [h_s, 0] For shallow nodes (depth <= h_s): only S-levels mapped to [depth, 0]
- Parameters:
bottom_depth – Positive bottom depth in meters
- Returns:
Array of depth values (negative, from deep to surface)
- classmethod read(filepath) SchismVgrid[source]
Read vgrid.in file. Supports two formats:
- Simple format (68 lines):
Line 0: nvrt kz h_s Lines 2-kz+1: Z-levels Remaining: S-levels
- LSC2 format (1.6GB, per-node):
Line 0: ivcor (=1) Line 1: nvrt Line 2: per-node kbp values … (per-node sigma levels)