Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Get from zero to a working osdf:// data read in a few minutes, then branch into the rest of the collection. This page uses ERA5 (NCAR GDEX dataset d633000) as a single running example so you can see every access pattern against the same data.

1. Why OSDF?

If accessing scientific data still means “download a giant archive, then analyze it locally,” OSDF is the alternative: it streams data over HTTPS to wherever your code runs — a laptop, a cloud VM, or an HPC node — so you read only the bytes you actually use. In this repository you access OSDF through PelicanFS, an FSSpec client, so xarray, intake, and pandas all work unchanged.

2. Set up — pick one path

Fastest (laptop, minimal deps). Just enough to run the first read below:

pip install pelicanfs xarray zarr dask aiohttp

Full repository (all notebooks). Clone and build the pinned environment:

git clone https://github.com/NCAR/osdf-examples.git
cd osdf-examples
conda env create -n osdf -f environment.yml
conda activate osdf
jupyter lab

One-click Binder. Coming soon — a launch-in-your-browser badge is in the works with the NCAR CIRRUS team. Until then, use one of the two paths above.

3. Your first OSDF read

Open ERA5 2-metre temperature straight from OSDF. The store is opened lazily, so this returns in seconds without downloading the full archive:

import xarray as xr

ds = xr.open_zarr("osdf:///ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr")
ds          # a lazy xarray.Dataset — metadata now, bytes on demand

Pull a single value to prove data really flows over OSDF:

ds["VAR_2T"].isel(time=0, latitude=360, longitude=720).values   # one point, streamed on demand

4. What just happened?

You asked a Pelican director for ncar/gdex/d633000/...; it pointed PelicanFS at a nearby cache (or the origin — the server that exposes GDEX data to the federation), and only the chunks you touched were transferred. You don’t have to manage any of this — but when you want control (a slow cache, a bad cache), the options in §5 give it to you. For the full concept primer, see the Introduction.

5. PelicanFS access patterns (with ERA5)

Zarr vs NetCDF

PelicanFS is format-agnostic — anything FSSpec can read works. Zarr stores open directly by URL:

ds = xr.open_zarr("osdf:///ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr")

NetCDF files are opened through a PelicanFS file handle (HDF5/NetCDF needs a seekable file object, so hand xr.open_dataset an open handle rather than a URL):

from pelicanfs.core import OSDFFileSystem

fs = OSDFFileSystem()
# One month of hourly ERA5 2-m temperature — a ~1.1 GB raw NetCDF file under d633000.
nc_path = "osdf:///ncar/gdex/d633000/e5.oper.an.sfc/202501/e5.oper.an.sfc.128_167_2t.ll025sc.2025010100_2025013123.nc"
with fs.open(nc_path, "rb") as f:
    ds_nc = xr.open_dataset(f, engine="h5netcdf")   # header now; variable data on demand

Opening reads only the file header, so it returns quickly despite the ~1 GB size — variable data streams in as you slice. For many NetCDF files at once, a kerchunk reference (below) is far more efficient than opening each one over HTTP.

Generic read (no options)

The everyday form — no extra arguments. The director rotates among healthy caches for you:

ds = xr.open_zarr("osdf:///ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr")

Reach for the options below only when the default path is slow or a cache is misbehaving.

Director-aware URL (without PelicanFS)

Because OSDF is just HTTPS, you can read without PelicanFS by handing the director URL straight to any HTTPS-aware reader — the director issues an HTTP redirect to a cache:

# Plain HTTPS via the director — no pelicanfs scheme, no cache/failover control.
ds = xr.open_zarr("https://osdf-director.osg-htc.org/ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr")

This is handy for a quick one-off, but you lose PelicanFS features — cache failover, direct_reads, and preferred_caches below. Prefer the osdf:/// scheme for real workflows.

direct_reads — go straight to the origin

Bypass the cache layer entirely and read from the origin. Useful when you’re close to the origin, or when a cache is serving bad data. For a Zarr open the option rides in storage_options:

ds = xr.open_zarr(
    "osdf:///ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr",
    storage_options={"direct_reads": True},
)

preferred_caches — pin the caches you trust

Try specific caches first. Keep the "+" sentinel so the director’s own caches remain as a fallback (without "+" there is no fallback — a single bad or mistyped host raises NoAvailableSource):

ds = xr.open_zarr(
    "osdf:///ncar/gdex/d633000/e5.oper.an.sfc.zarr/e5.oper.an.sfc.2t.zarr",
    storage_options={"preferred_caches": ["https://a-good-cache.example.org:8443", "+"]},
)

kerchunk references (virtual Zarr over NetCDF)

kerchunk exposes many NetCDF files as one virtual Zarr store via a small reference file — and both the reference and the data it points to can be read over OSDF. Cache-control options for the data chunks ride in remote_options here (not storage_options, as they would for a plain Zarr open):

ds = xr.open_dataset(
    "osdf:///ncar/gdex/d640000/kerchunk/anl_surf125-remote-osdf.parq",  # the reference, over OSDF
    engine="kerchunk",
    storage_options={
        "remote_protocol": "osdf",                   # the data targets are OSDF too
        "remote_options": {"direct_reads": True},    # cache control for the data chunks
    },                                               # (or {"preferred_caches": [...]}, or omit)
)

This is the JRA-3Q surface reference behind jja_heatindex.ipynb — a complete, working kerchunk-over-OSDF workflow. ERA5 references live under osdf:///ncar/gdex/d633000/kerchunk/... (e.g. the meanflux/ collection) and open exactly the same way.

6. Common errors and fixes

SymptomLikely causeFix
FileNotFoundError / nothing resolvesosdf://pathtwo slashesuse osdf:///path (three)
NoAvailableSourceno working cache — a pinned/mistyped host, or preferred_caches with no "+" fallbackcheck the host; add "+"; drop bad entries
ReferenceNotReachable / ConnectionErrortoo many concurrent range reads (kerchunk fan-out) overwhelm a cachereduce concurrency (fewer Dask workers / lower zarr async.concurrency), retry, or direct_reads
zlib: incorrect header check / ContentLengthError, every retrya cache serving corrupt bytesdirect_reads=True (origin) or pin a good cache via preferred_caches
kerchunk data won’t materialize / silently re-reads.load()/.compute() return lazy Dask on this stackpull .values / np.asarray(...) to force true NumPy
MergeError: conflicting values for 'utc_date'merging ERA5 surface stores with a differing scalar coordsub-select variables before merge, or xr.merge(..., compat="override")

7. Find your next notebook

8. Go deeper