Data Service

Package and test-result data is consolidated in a single on-disk store, written by exactly one producer and read by everything else.

The store

The store root is in $SKA/data/skare3/skare3_data/data (configured in skare3_tools via CONFIG["data_dir"]). It holds:

  • manifest.json — schema version, generation time, producer identity, and the excluded repositories,

  • repository_status.json — operator-edited input mapping owner/repo to "deprecated" or "ignored" (both excluded from the store); seeded once by the first refresh, never overwritten,

  • packages.json — every repository’s record, plus the channel and metapackage versions; this is what the dashboards consume,

  • package_list.json — the package universe (skare3 recipes + organization repositories), unfiltered,

  • test_results.json — the digested latest regression-test results,

  • test_logs/ — all test-results runs (see test_results),

  • meta/ — producer bookkeeping (change-detection state, lock).

Every file is written atomically, so readers (including rsync’d copies) never see a half-written file.

One copy of each record

packages.json holds each repository’s record exactly once: there is no per-repository tree beside it that could drift out of sync. This file is incrementally updated by the skare3-refresh script.

There are two different recorded versions in this file: schema_version and record_version:

  • schema_version describes the store’s layout, specifying which files exist and how a reader finds things in them. A reader of another version declines the store rather than half-reading it.

  • record_version (with record_options) describes the shape of one repository record. This is what decides whether the producer can reuse the already existing record.

Keeping them apart matters: a layout change must not cost a refetch of every repository, which would be thousands of GraphQL queries for records that are not invalid.

Package Deployment Status

There are five fields related to the package’s deployment status:

  • master_version is the latest version in the ska3-masters channel.

  • flight is the version pinned by the ska3-flight metapackage.

  • matlab is the version pinned by the ska3-matlab metapackage.

  • aca is the version pinned by the ska3-aca metapackage.

  • test_version is the version used in the latest regression run.

  • test_status is the latest regression run result (pass or fail).

These fields are not directly related to repository activity, so every refresh recomputes all of them for every repository — including the ones whose details were not refetched. A master_version that does not correspond to the master branch is a CI failure to investigate.

The consolidated package-data store.

One producer (skare3-refresh) writes these files; everything else reads them. The store root is CONFIG["data_dir"] — $SKA/data/skare3/skare3_data/data on synced hosts — which reaches every machine through the existing $SKA/data rsync. Layout:

<data_dir>/
├── manifest.json          # schema_version, generated, producer, excluded
├── repository_status.json # operator-edited: {owner/repo: "deprecated"|"ignored"}
├── packages.json          # the per-repository records (the only copy) + channel state
├── package_list.json      # the package universe (pkg_defs + org repos), unfiltered
├── test_results.json      # pre-digested latest test results
├── test_logs/             # test-results runs (see test_results.py)
└── meta/
    ├── state.json         # producer bookkeeping (last-updated timestamps)
    └── refresh.lock

packages.json holds each repository’s record exactly once: there is no separate per-repository tree to drift out of sync with it. Single-repository reads pick the entry out of the aggregate (see StoreReader).

Every file is written atomically (temp file + os.replace), so readers on rsync’d copies never see a half-written file. Readers never take the lock.

repository_status.json is the one file the producer reads rather than writes: repositories listed there (with either status) are excluded from the store. The first refresh seeds it if absent; after that it is never overwritten.

The recipes

The package universe comes from the skare3 recipes (pkg_defs/*/meta.yaml) plus the organizations’ repository listings. The recipes are fetched as a tarball of the default branch into a temporary directory, once per run, and nothing is left behind: there is deliberately no checkout in the data directory.

If the fetch fails, the store’s own package_list.json is the parsed product of an earlier fetch and stands in for the recipes. The error is then reported in a summary. If the fetch fails and there is no package_list.json, the run aborts.

The producer: skare3-refresh

The single writer of the package-data store.

The command-line entry point is skare3_tools.scripts.refresh (skare3-refresh).

One run brings the store (see skare3_tools.packages.store) up to date:

  1. refresh the package list (the skare3 recipes, fetched to a temporary directory, plus the org repositories) into package_list.json, and take the working universe from it, excluding repositories listed in repository_status.json at the store root — an operator-edited file that refresh seeds once if missing and never overwrites,

  2. snapshot the conda channels once and resolve the four metapackages (ska3-aca/flight/matlab/perl) — failing loudly if any can’t be resolved,

  3. detect changed repositories with one batched GraphQL query and fetch detail only for those, carrying the rest over from the previous packages.json (the aggregate is the incremental cache as well as the output — there is no second copy of a repository’s record anywhere),

  4. rebuild packages.json (always — metapackage pins and channel versions can change without any repository push), digest the latest test results, and advance meta/state.json last, so an interrupted run only causes a refetch.

If there is no readable test run, the tested versions already in the store are kept rather than blanked, and the run is reported as a failure: an aggregate claiming nothing was tested is indistinguishable from the truth once written.

If the conda channels or GitHub cannot be reached at all, the run aborts with RefreshError before writing anything: the store keeps the data it has.

Authentication is entirely the github wrappers’ business: a personal token (GITHUB_API_TOKEN/GITHUB_TOKEN) or, when SKARE3_GITHUB_APP_KEY is set, per-organization App-77359 installation tokens minted transparently per request (see skare3_tools.github.app_auth).

Publishing: skare3-dashboard-update

The hourly production job reduces to skare3-refresh followed by rendering, and “publishing” amounts to copying the JSON files into the public directory served over HTTP.

Publish the data store for HTTP readers.

skare3-dashboard-update is the hourly production entry point: refresh the store, then copy the published subset into the public dashboard directory — packages.json and test_results.json (fetched by the React dashboard), plus package_list.json and repository_status.json (the DataClient HTTP tier).

The store’s manifest.json and meta/ are never published: the manifest name collides with the React app’s own manifest.json at the public root (and the HTTP tier does not read it), and meta/ is producer bookkeeping.

Files are copied via a temp name + os.replace so HTTP readers see the old or the new content, never a partial file.

Reading the data: DataClient

class skare3_tools.packages.DataClient(source='auto', data_dir=None, url=None)[source]

Read package data from the best available source.

Parameters:
  • source – “auto”, “local”, “http” or “github”.

  • data_dir – local store directory (default: the configured root).

  • url – published store URL (default: CONFIG[“store_url”]).

generated()[source]

When the data was produced (ISO string), or “” if it does not say.

package_list()[source]

The package universe, as get_package_list.

packages()[source]

The aggregate: every repository’s record, plus the channel versions.

repository_info(owner_repo)[source]

Detailed information for one repository.

The aggregate holds the only copy of each record, so this picks the entry out of it – identically for every source.

sources()[source]

The sources to try, in order, for this client’s configured source.

test_results()[source]

The digested latest test results (test_results.json).

Falling back is decided per file, not once per client: a store or a published location can be missing one file and be authoritative for the rest — which is exactly what a store written by an older layout looks like. One absent file must not send every other read to Github.

The public functions of skare3_tools.packages (get_repository_info, get_repositories_info and get_package_list) read the store through this client, so they need no Github token. They query Github only when asked for something the store does not hold — a different since, say — or with update=True. What the store’s records were produced with is recorded in the aggregate itself, and reported by record_options.

The dashboard views (skare3-dashboard, skare3-test-dashboard) are plain renderers of this data: they read it through this client, and do not query Github themselves.