Downloads tables from the POLIS OData service to a local cache, with resumable batches and optional parallel downloads by year. The downloader accounts for these API behaviours:
Date filters are year-aligned only. POLIS only honours
field le YYYY-12-31bounds; any sub-yearlereturns 0 rows. The function alignsmin_date/max_dateto year boundaries before building any filter.No
@odata.nextLinkand$skipis rejected. POLIS caps$topat 2000 and refuses to paginate via the OData standard mechanism. The only way to walk past 2000 rows is Id-range pagination:$orderby=Id&$top=2000&$filter=... and Id gt <last>.The clinical date columns (
CaseDate,VirusDate, ...) have NULL coverage for historical records. The function filters on thedate_fieldin polis_tables_mapping: a revision date (LastUpdateDate/UpdatedDate) or an event date (Start/PublishDate).
Usage
get_polis_data(
tables = NULL,
min_date = "2000-01-01",
max_date = Sys.Date(),
region = "Global",
country_code = NULL,
polis_folder = tools::R_user_dir("polished", which = "cache"),
output_format = c("rds", "rda", "csv", "parquet", "qs2"),
workers = 1L,
auto_refetch = TRUE,
verify_years = 3L,
log_file = NULL,
keep_archives = 0L,
force = FALSE,
prune_parts = TRUE,
polis_api_key = Sys.getenv("POLIS_API_KEY"),
quiet = FALSE,
reference_refresh_days = 1
)Arguments
- tables
Optional character vector of table names (see polis_tables_mapping for the supported set, e.g.
"case","virus","population").NULL(default) downloads every table in the catalogue. Unknown names abort with a list of valid names. Thepopulationreference table has no update date, somin_date/max_date/regiondo not apply to it – it is pulled whole.- min_date
Earliest date to fetch. Defaults to
"2000-01-01". POLIS rejects sub-year date ranges, so this is aligned to January 1 of its year.- max_date
Latest date to fetch. Defaults to
Sys.Date(). Aligned to December 31 of its year.- region
WHO region filter (
"Global"(default),"AFRO","AMRO","EMRO","EURO","SEARO","WPRO").- country_code
Optional ISO3 country code (e.g.
"NGA"). Adds anand CountryISO3Code eq '<code>'clause. DefaultNULL(no country filter).- polis_folder
Root folder for cached data. Files land under
<polis_folder>/. Defaulttools::R_user_dir("polished", which = "cache")– the standard per-user cache location, persistent across sessions so incremental updates reuse the same cache. Pass an explicit path to keep data alongside a project.- output_format
Output format. One of
"rds"(default),"rda","csv","parquet","qs2"."parquet"requires thearrowpackage;"qs2"requires theqs2package.- workers
Number of parallel workers.
1(default) runs a sequential loop with a live per-batch progress bar.> 1opts into a PSOCK cluster that dispatches one year per worker; pass e.g.parallel::detectCores() - 1Lto use most cores.- auto_refetch
If
TRUE(default), verify fresh downloads and refresh completed snapshots as described above.FALSEtrusts a completed snapshot with matching scope and skips post-download verification.- verify_years
Number of most recent calendar years covered by the missing-Id check on a fresh pull, counted back from
max_dateand clamped tomin_date. Default3L;NULLcovers the full requested range. Revision-based reconciliation always covers the full scope. Ignored whenauto_refetch = FALSEand for reference tables with no update date.- log_file
Optional path to a per-batch log file (
.rds). DefaultNULL.- keep_archives
When
> 0, on each save also writes a timestamped copy underarchive/and prunes older copies. Default0(no archive).- force
If
TRUE, discard resume parts and fetch a fresh snapshot. The previous canonical file remains until replacement succeeds.- prune_parts
If
TRUE(default), delete the per-year resume cache after publishing the complete snapshot.FALSEretains completed parts for inspection; they are discarded when the snapshot's revisions change.- polis_api_key
API key. Defaults to
Sys.getenv("POLIS_API_KEY").- quiet
Suppress headers, progress bars, and the info alert. Default
FALSE.- reference_refresh_days
Maximum age in days of a completed reference table without an update timestamp (currently
population). Default1;0refreshes on every call. Ignored whenauto_refetch = FALSE.
Value
NULL, invisibly. Each selected table is written to
<polis_folder>/<raw_stem>.<ext>. Pages and year partitions are loaded
during downloading; the final merge loads the complete table into memory.
Read a saved table with readRDS(file.path(polis_folder, "raw_im.rds")).
Details
Cache layout and recovery. Pages are committed under
<polis_folder>/.parts/<stem>/year_YYYY.<ext>.pages/, then compacted once
per year. A journal records the last committed page. Interrupted pulls
resume at that page's maximum Id; an uncommitted page is fetched again.
After all years finish, the verified table replaces
<polis_folder>/<stem>.<ext> atomically. The stem is the table's raw_*
name (e.g. case is written as raw_afp; see polis_tables_mapping).
Merging requires memory for the complete table.
Scope manifests under .manifests/ record the effective country, region,
year range and format. Changing those filters starts a fresh pull while
preserving the previous canonical file until the replacement succeeds.
Legacy files are renamed to their raw_* names, but caches without a scope
manifest must be downloaded once again because their filters are unknown.
prune_parts = TRUE removes completed parts; unchanged snapshots can be
reused without splitting or reading the full canonical file.
Parallelism. Each calendar year between min_date and max_date
is an independent Id-range walk. With workers > 1, years are
dispatched across a parallel::makePSOCKcluster() cluster (the same
transport on Windows, macOS, and Linux, so any user can opt in); a
live cli progress bar polls the part files between socket reads so
you see rows accumulate across workers in real time. With
workers = 1, a single sequential loop drives the bar per 2K-row
batch. PSOCK workers need polished installed in their library
path – devtools::load_all() is not enough.
Freshness. For tables with LastUpdateDate or UpdatedDate, completed
snapshots are checked against the full scoped list of Ids and revision
timestamps. Changed and new rows are fetched selectively; deleted rows are
removed. Equal row counts never establish freshness. This relies on the
service advancing its revision timestamp when a record changes.
Tables filtered by event dates (Start or PublishDate) are fetched again
because those dates cannot establish whether a row was edited.
Reference tables with no update date use reference_refresh_days instead.
force = TRUE always starts a full pull. auto_refetch = FALSE explicitly
trusts completed snapshots until forced or their scope changes.
Fresh pulls also check for missing Ids over verify_years and refetch them.
Revision-based reconciliation covers the full scope, including after an
interrupted pull resumes. Verification failures retain checkpoints and
leave the previous canonical file in place for a later retry.
Resilience. Read timeouts are retried inside each request, and a year
whose parallel worker still fails is requeued up to three times.
POLIS_TIMEOUT_SECONDS overrides the 120-second default.
See also
polis_tables_mapping for the table catalogue.
Examples
if (FALSE) { # \dontrun{
# Pull one table into the default per-user cache
get_polis_data(tables = "im")
# Read it back from disk when you need it
cache <- tools::R_user_dir("polished", which = "cache")
im <- readRDS(file.path(cache, "raw_im.rds"))
# The whole catalogue in parallel into a project folder
get_polis_data(
polis_folder = "data/polis",
workers = parallel::detectCores() - 1L
)
} # }
