Skip to contents

Downloads tables from the POLIS OData service to a local cache, with resumable batches and optional parallel downloads by year. The downloader accounts for these API behaviours:

  • Date filters are year-aligned only. POLIS only honours field le YYYY-12-31 bounds; any sub-year le returns 0 rows. The function aligns min_date / max_date to year boundaries before building any filter.

  • No @odata.nextLink and $skip is rejected. POLIS caps $top at 2000 and refuses to paginate via the OData standard mechanism. The only way to walk past 2000 rows is Id-range pagination: $orderby=Id&$top=2000&$filter=... and Id gt <last>.

  • The clinical date columns (CaseDate, VirusDate, ...) have NULL coverage for historical records. The function filters on the date_field in polis_tables_mapping: a revision date (LastUpdateDate / UpdatedDate) or an event date (Start / PublishDate).

Usage

get_polis_data(
  tables = NULL,
  min_date = "2000-01-01",
  max_date = Sys.Date(),
  region = "Global",
  country_code = NULL,
  polis_folder = tools::R_user_dir("polished", which = "cache"),
  output_format = c("rds", "rda", "csv", "parquet", "qs2"),
  workers = 1L,
  auto_refetch = TRUE,
  verify_years = 3L,
  log_file = NULL,
  keep_archives = 0L,
  force = FALSE,
  prune_parts = TRUE,
  polis_api_key = Sys.getenv("POLIS_API_KEY"),
  quiet = FALSE,
  reference_refresh_days = 1
)

Arguments

tables

Optional character vector of table names (see polis_tables_mapping for the supported set, e.g. "case", "virus", "population"). NULL (default) downloads every table in the catalogue. Unknown names abort with a list of valid names. The population reference table has no update date, so min_date/max_date/region do not apply to it – it is pulled whole.

min_date

Earliest date to fetch. Defaults to "2000-01-01". POLIS rejects sub-year date ranges, so this is aligned to January 1 of its year.

max_date

Latest date to fetch. Defaults to Sys.Date(). Aligned to December 31 of its year.

region

WHO region filter ("Global" (default), "AFRO", "AMRO", "EMRO", "EURO", "SEARO", "WPRO").

country_code

Optional ISO3 country code (e.g. "NGA"). Adds an and CountryISO3Code eq '<code>' clause. Default NULL (no country filter).

polis_folder

Root folder for cached data. Files land under <polis_folder>/. Default tools::R_user_dir("polished", which = "cache") – the standard per-user cache location, persistent across sessions so incremental updates reuse the same cache. Pass an explicit path to keep data alongside a project.

output_format

Output format. One of "rds" (default), "rda", "csv", "parquet", "qs2". "parquet" requires the arrow package; "qs2" requires the qs2 package.

workers

Number of parallel workers. 1 (default) runs a sequential loop with a live per-batch progress bar. > 1 opts into a PSOCK cluster that dispatches one year per worker; pass e.g. parallel::detectCores() - 1L to use most cores.

auto_refetch

If TRUE (default), verify fresh downloads and refresh completed snapshots as described above. FALSE trusts a completed snapshot with matching scope and skips post-download verification.

verify_years

Number of most recent calendar years covered by the missing-Id check on a fresh pull, counted back from max_date and clamped to min_date. Default 3L; NULL covers the full requested range. Revision-based reconciliation always covers the full scope. Ignored when auto_refetch = FALSE and for reference tables with no update date.

log_file

Optional path to a per-batch log file (.rds). Default NULL.

keep_archives

When > 0, on each save also writes a timestamped copy under archive/ and prunes older copies. Default 0 (no archive).

force

If TRUE, discard resume parts and fetch a fresh snapshot. The previous canonical file remains until replacement succeeds.

prune_parts

If TRUE (default), delete the per-year resume cache after publishing the complete snapshot. FALSE retains completed parts for inspection; they are discarded when the snapshot's revisions change.

polis_api_key

API key. Defaults to Sys.getenv("POLIS_API_KEY").

quiet

Suppress headers, progress bars, and the info alert. Default FALSE.

reference_refresh_days

Maximum age in days of a completed reference table without an update timestamp (currently population). Default 1; 0 refreshes on every call. Ignored when auto_refetch = FALSE.

Value

NULL, invisibly. Each selected table is written to <polis_folder>/<raw_stem>.<ext>. Pages and year partitions are loaded during downloading; the final merge loads the complete table into memory. Read a saved table with readRDS(file.path(polis_folder, "raw_im.rds")).

Details

Cache layout and recovery. Pages are committed under <polis_folder>/.parts/<stem>/year_YYYY.<ext>.pages/, then compacted once per year. A journal records the last committed page. Interrupted pulls resume at that page's maximum Id; an uncommitted page is fetched again. After all years finish, the verified table replaces <polis_folder>/<stem>.<ext> atomically. The stem is the table's raw_* name (e.g. case is written as raw_afp; see polis_tables_mapping). Merging requires memory for the complete table.

Scope manifests under .manifests/ record the effective country, region, year range and format. Changing those filters starts a fresh pull while preserving the previous canonical file until the replacement succeeds. Legacy files are renamed to their raw_* names, but caches without a scope manifest must be downloaded once again because their filters are unknown. prune_parts = TRUE removes completed parts; unchanged snapshots can be reused without splitting or reading the full canonical file.

Parallelism. Each calendar year between min_date and max_date is an independent Id-range walk. With workers > 1, years are dispatched across a parallel::makePSOCKcluster() cluster (the same transport on Windows, macOS, and Linux, so any user can opt in); a live cli progress bar polls the part files between socket reads so you see rows accumulate across workers in real time. With workers = 1, a single sequential loop drives the bar per 2K-row batch. PSOCK workers need polished installed in their library path – devtools::load_all() is not enough.

Freshness. For tables with LastUpdateDate or UpdatedDate, completed snapshots are checked against the full scoped list of Ids and revision timestamps. Changed and new rows are fetched selectively; deleted rows are removed. Equal row counts never establish freshness. This relies on the service advancing its revision timestamp when a record changes. Tables filtered by event dates (Start or PublishDate) are fetched again because those dates cannot establish whether a row was edited. Reference tables with no update date use reference_refresh_days instead. force = TRUE always starts a full pull. auto_refetch = FALSE explicitly trusts completed snapshots until forced or their scope changes.

Fresh pulls also check for missing Ids over verify_years and refetch them. Revision-based reconciliation covers the full scope, including after an interrupted pull resumes. Verification failures retain checkpoints and leave the previous canonical file in place for a later retry.

Resilience. Read timeouts are retried inside each request, and a year whose parallel worker still fails is requeued up to three times. POLIS_TIMEOUT_SECONDS overrides the 120-second default.

See also

polis_tables_mapping for the table catalogue.

Examples

if (FALSE) { # \dontrun{
# Pull one table into the default per-user cache
get_polis_data(tables = "im")

# Read it back from disk when you need it
cache <- tools::R_user_dir("polished", which = "cache")
im <- readRDS(file.path(cache, "raw_im.rds"))

# The whole catalogue in parallel into a project folder
get_polis_data(
  polis_folder = "data/polis",
  workers = parallel::detectCores() - 1L
)
} # }