Data Service¶
Package and test-result data is consolidated in a single on-disk store, written by exactly one producer and read by everything else.
The store¶
The store root is in $SKA/data/skare3/skare3_data/data (configured in skare3_tools via
CONFIG["data_dir"]). It holds:
manifest.json— schema version, generation time, producer identity, and the excluded repositories,repository_status.json— operator-edited input mappingowner/repoto"deprecated"or"ignored"(both excluded from the store); seeded once by the first refresh, never overwritten,packages.json— every repository’s record, plus the channel and metapackage versions; this is what the dashboards consume,package_list.json— the package universe (skare3 recipes + organization repositories), unfiltered,test_results.json— the digested latest regression-test results,test_logs/— all test-results runs (see test_results),meta/— producer bookkeeping (change-detection state, lock).
Every file is written atomically, so readers (including rsync’d copies) never see a half-written file.
One copy of each record¶
packages.json holds each repository’s record exactly once: there is no
per-repository tree beside it that could drift out of sync. This file is incrementally
updated by the skare3-refresh script.
There are two different recorded versions in this file: schema_version and record_version:
schema_versiondescribes the store’s layout, specifying which files exist and how a reader finds things in them. A reader of another version declines the store rather than half-reading it.record_version(withrecord_options) describes the shape of one repository record. This is what decides whether the producer can reuse the already existing record.
Keeping them apart matters: a layout change must not cost a refetch of every repository, which would be thousands of GraphQL queries for records that are not invalid.
Package Deployment Status¶
There are five fields related to the package’s deployment status:
master_versionis the latest version in theska3-masterschannel.flightis the version pinned by theska3-flightmetapackage.matlabis the version pinned by theska3-matlabmetapackage.acais the version pinned by theska3-acametapackage.test_versionis the version used in the latest regression run.test_statusis the latest regression run result (pass or fail).
These fields are not directly related to repository activity, so
every refresh recomputes all of them for every repository — including the ones
whose details were not refetched. A master_version that does not correspond
to the master branch is a CI failure to investigate.
The consolidated package-data store.
One producer (skare3-refresh) writes these files; everything else reads
them. The store root is CONFIG["data_dir"] — $SKA/data/skare3/skare3_data/data
on synced hosts — which reaches every machine through the existing $SKA/data
rsync. Layout:
<data_dir>/
├── manifest.json # schema_version, generated, producer, excluded
├── repository_status.json # operator-edited: {owner/repo: "deprecated"|"ignored"}
├── packages.json # the per-repository records (the only copy) + channel state
├── package_list.json # the package universe (pkg_defs + org repos), unfiltered
├── test_results.json # pre-digested latest test results
├── test_logs/ # test-results runs (see test_results.py)
└── meta/
├── state.json # producer bookkeeping (last-updated timestamps)
└── refresh.lock
packages.json holds each repository’s record exactly once: there is no
separate per-repository tree to drift out of sync with it. Single-repository
reads pick the entry out of the aggregate (see StoreReader).
Every file is written atomically (temp file + os.replace), so readers on
rsync’d copies never see a half-written file. Readers never take the lock.
repository_status.json is the one file the producer reads rather than
writes: repositories listed there (with either status) are excluded from the
store. The first refresh seeds it if absent; after that it is never overwritten.
The recipes¶
The package universe comes from the skare3 recipes (pkg_defs/*/meta.yaml)
plus the organizations’ repository listings. The recipes are fetched as a
tarball of the default branch into a temporary directory, once per run, and
nothing is left behind: there is deliberately no checkout in the data
directory.
If the fetch fails, the store’s own package_list.json is the parsed product
of an earlier fetch and stands in for the recipes. The error is then reported in a summary.
If the fetch fails and there is no package_list.json, the run aborts.
The producer: skare3-refresh¶
The single writer of the package-data store.
The command-line entry point is skare3_tools.scripts.refresh
(skare3-refresh).
One run brings the store (see skare3_tools.packages.store) up to date:
refresh the package list (the skare3 recipes, fetched to a temporary directory, plus the org repositories) into
package_list.json, and take the working universe from it, excluding repositories listed inrepository_status.jsonat the store root — an operator-edited file that refresh seeds once if missing and never overwrites,snapshot the conda channels once and resolve the four metapackages (ska3-aca/flight/matlab/perl) — failing loudly if any can’t be resolved,
detect changed repositories with one batched GraphQL query and fetch detail only for those, carrying the rest over from the previous
packages.json(the aggregate is the incremental cache as well as the output — there is no second copy of a repository’s record anywhere),rebuild
packages.json(always — metapackage pins and channel versions can change without any repository push), digest the latest test results, and advancemeta/state.jsonlast, so an interrupted run only causes a refetch.
If there is no readable test run, the tested versions already in the store are kept rather than blanked, and the run is reported as a failure: an aggregate claiming nothing was tested is indistinguishable from the truth once written.
If the conda channels or GitHub cannot be reached at all, the run aborts with
RefreshError before writing anything: the store keeps the data it has.
Authentication is entirely the github wrappers’ business: a personal token
(GITHUB_API_TOKEN/GITHUB_TOKEN) or, when SKARE3_GITHUB_APP_KEY is
set, per-organization App-77359 installation tokens minted transparently per
request (see skare3_tools.github.app_auth).
Publishing: skare3-dashboard-update¶
The hourly production job reduces to skare3-refresh followed by rendering, and “publishing”
amounts to copying the JSON files into the public directory served over HTTP.
Publish the data store for HTTP readers.
skare3-dashboard-update is the hourly production entry point: refresh the
store, then copy the published subset into the public dashboard directory —
packages.json and test_results.json (fetched by the React dashboard),
plus package_list.json and repository_status.json (the DataClient HTTP
tier).
The store’s manifest.json and meta/ are never published: the manifest
name collides with the React app’s own manifest.json at the public root
(and the HTTP tier does not read it), and meta/ is producer bookkeeping.
Files are copied via a temp name + os.replace so HTTP readers see the old
or the new content, never a partial file.
Reading the data: DataClient¶
- class skare3_tools.packages.DataClient(source='auto', data_dir=None, url=None)[source]¶
Read package data from the best available source.
- Parameters:
source – “auto”, “local”, “http” or “github”.
data_dir – local store directory (default: the configured root).
url – published store URL (default: CONFIG[“store_url”]).
- package_list()[source]¶
The package universe, as
get_package_list.
Falling back is decided per file, not once per client: a store or a published location can be missing one file and be authoritative for the rest — which is exactly what a store written by an older layout looks like. One absent file must not send every other read to Github.
The public functions of skare3_tools.packages
(get_repository_info,
get_repositories_info and
get_package_list) read the store through this
client, so they need no Github token. They query Github only when asked for
something the store does not hold — a different since, say — or with
update=True. What the store’s records were produced with is recorded in
the aggregate itself, and reported by
record_options.
The dashboard views (skare3-dashboard, skare3-test-dashboard) are plain renderers of this
data: they read it through this client, and do not query Github themselves.