Sys.setenv(POLIS_API_KEY = "your-key") # or set it in .Renviron
# Pull the immunization (`im`) table; the data is written to the per-user
# cache and nothing is returned
polished::get_polis_data(tables = "im")
# Read the table back from disk when you need it (written under its raw_* name)
cache <- tools::R_user_dir("polished", which = "cache")
im <- readRDS(file.path(cache, "raw_im.rds"))Note
The download calls in this vignette talk to the live POLIS service and are shown unevaluated — they need a valid
POLIS_API_KEYand network access. The only chunk that runs at build time is the table catalogue, which is static package data. Copy the calls into a session with a key set to try them.
What get_polis_data() is for
polished::get_polis_data() downloads tables from the POLIS OData API, caches them on disk and resumes interrupted downloads from saved checkpoints.
The downloader accounts for these API behaviours:
-
Date filters are year-aligned only. POLIS honours a
field le YYYY-12-31bound but returns zero rows for any sub-yearle.get_polis_data()aligns yourmin_date/max_dateto year boundaries before building a filter, so a mid-year range never silently empties the result. -
$skipis rejected and there is no@odata.nextLink. POLIS caps$topat 2000 rows and refuses the standard OData paging mechanism. The only way past 2000 rows is Id-range pagination —…&$orderby=Id&$top=2000&$filter=… and Id gt <last>— which the function does for you, one page at a time. -
The clinical date columns are sparsely populated. Columns like
CaseDateandVirusDateareNULLfor many historical records, so filtering on them can exclude those rows. The function uses thedate_fieldlisted inpolis_tables_mapping: a revision date (LastUpdateDateorUpdatedDate) or an event date (StartorPublishDate), depending on the table.
Quick start
Set your key once, then pull a table. By default nothing is returned — the data lands on disk and you read it when you need it.
The table catalogue
tables accepts any of the names in polis_tables_mapping, the static catalogue the function ships with. Passing tables = NULL (the default) downloads them all. This is the one chunk in the vignette that runs at build time, because it is just package data:
polished::polis_tables_mapping
#> table_name endpoint date_field
#> 1 virus Virus UpdatedDate
#> 2 case Case LastUpdateDate
#> 3 human_specimen LabSpecimen LastUpdateDate
#> 4 environmental_sample EnvSample LastUpdateDate
#> 5 activity Activity LastUpdateDate
#> 6 sub_activity SubActivity UpdatedDate
#> 7 lqas Lqas Start
#> 8 im Im PublishDate
#> 9 historized_synonyms HistorizedSynonyms LastUpdateDate
#> 10 historized_geoplace_names HistorizedGeoplaceNames LastUpdateDate
#> 11 population Population <NA>
#> file_stem
#> 1 raw_virus
#> 2 raw_afp
#> 3 raw_hum_spec
#> 4 raw_es
#> 5 raw_activity
#> 6 raw_sub_activity
#> 7 raw_lqas
#> 8 raw_im
#> 9 raw_historized_synonyms
#> 10 raw_historized_geoplace_names
#> 11 raw_populationEach row records the short table_name you pass to tables =, the OData endpoint it maps to, the date_field used for both filtering and the duplicate-resolution tiebreak, and the file_stem the table is written under on disk. That stem is the raw_* name the cleaning pipeline reads (e.g. case is saved as raw_afp), so the download and cleaning halves share one naming convention. An unknown name aborts with the list of valid ones, so a typo fails fast rather than fetching nothing.
Choosing what to fetch
Four arguments narrow the pull:
| Argument | Effect |
|---|---|
min_date / max_date
|
the date range, aligned to whole years (default 2000-01-01 to today) |
region |
a WHO region filter ("Global", "AFRO", "EMRO", …) |
country_code |
an ISO3 code (e.g. "NGA") added as an exact-match clause |
# Nigeria only, 2018 through last year, AFRO region
polished::get_polis_data(
tables = "case",
min_date = "2018-01-01",
max_date = "2024-12-31",
region = "AFRO",
country_code = "NGA"
)Because date bounds are year-aligned, min_date = "2018-06-30" fetches from 2018-01-01 regardless — the function never returns fewer rows than the year covering your bound.
Where the data lives, and how resume works
Each batch is written as an immutable page and committed in a journal. Once a year finishes, its pages are compacted into one part. The final merge writes the canonical file under the table’s raw_* stem:
<polis_folder>/
├── raw_afp.rds
├── .manifests/raw_afp.rds.rds # scope, revisions and completion state
└── .parts/raw_afp/
├── year_2023.rds
├── year_2023.meta.rds # progress without rereading the data
└── year_2024.rds.pages/
├── state.rds # committed pages and last Id
└── page_00000001.rds
An interrupted run resumes after the last committed page. A page whose journal commit failed is fetched again. The final merge needs memory for the complete table. A failed refresh leaves the previous canonical file in place.
Files with older names such as case.rds are renamed to raw_afp.rds. Legacy caches without scope metadata require one fresh download because the country, region and date filters used to create them are unknown.
polis_folder defaults to tools::R_user_dir("polished", which = "cache"), the standard per-user cache location, so incremental updates persist across sessions without you choosing a path. Pass an explicit folder to keep data alongside a project instead:
polished::get_polis_data(tables = "im", polis_folder = "data/polis")Incremental updates
Re-running a call checks the saved scope before reusing its data. Country, region or year-range changes start a fresh pull. For tables with LastUpdateDate or UpdatedDate, the downloader compares the full scoped Id and revision list, fetches changed rows, and removes deleted records. This assumes the service advances its revision timestamp on edits. Equal row counts do not establish freshness.
An unchanged snapshot needs no canonical read or partition rebuilding. Tables filtered by event dates (lqas and im) are fetched again because those dates do not identify edits. Population snapshots have a one-day expiry by default; use reference_refresh_days = 0 to refresh them on every call.
force = TRUE starts a full pull while retaining the old canonical until its replacement succeeds. auto_refetch = FALSE explicitly trusts completed snapshots with matching scope until forced.
By default, prune_parts = TRUE removes completed parts. Use FALSE to retain them for inspection. Revision changes discard retained parts so they cannot restore stale records on a later run.
Going faster with parallel workers
Each calendar year is an independent walk, so years can be fetched concurrently. workers > 1 dispatches them across a parallel::makePSOCKcluster() — the same transport on Windows, macOS, and Linux — while a single live progress bar polls the part files so you watch rows accumulate across workers in real time:
polished::get_polis_data(
tables = "virus",
workers = parallel::detectCores() - 1L
)Important
PSOCK workers start fresh R sessions and load
polishedwithlibrary(), so the package must be installed for parallel mode — adevtools::load_all()session is not enough. Withworkers = 1L(the default) a single sequential loop drives the bar per batch and has no such requirement.
What you get back
get_polis_data() writes each table to disk and returns NULL invisibly. Read the tables you need with the appropriate file reader. The download’s final merge still requires memory for the complete table.
Each table lands at <polis_folder>/<file_stem>.<ext> (the raw_* name from the catalogue). Read one back with the matching reader for your output_format:
# default per-user cache location
cache <- tools::R_user_dir("polished", which = "cache")
case <- readRDS(file.path(cache, "raw_afp.rds")) # `case` is saved as raw_afp
# or, for other output formats:
# arrow::read_parquet(file.path(cache, "raw_afp.parquet"))Assigning the call stores NULL, not the output paths. Construct paths from polis_folder, the table’s file_stem and output_format as shown above.
Output formats
output_format controls how the canonical file is written: "rds" (default), "rda", "csv", "parquet", or "qs2". The parquet and qs2 formats need the arrow and qs2 packages respectively.
polished::get_polis_data(tables = "case", output_format = "parquet")Completeness verification
POLIS occasionally truncates a query under load and returns a partial page even when more rows exist. With auto_refetch = TRUE (the default), after each table finishes the function checks the most recent verify_years calendar years (default 3L). Set verify_years = NULL to check the full date range. It:
- uses the metadata sidecars to detect a gap cheaply (row-count and Id-range mismatch against POLIS’s reported
@odata.count); - only if a gap is found, issues a lightweight
$select=Idprobe to list the canonical Id set; and - refetches any missing Ids via
Id in (...)chunks and merges them in, de-duplicating byIdkeeping the latest update date.
Set auto_refetch = FALSE to reuse completed snapshots with matching scope without freshness or completeness checks. Revision-based reconciliation, when enabled, covers the full requested scope independently of verify_years.
Other options worth knowing
| Argument | What it does |
|---|---|
keep_archives |
when > 0, also writes a timestamped copy under archive/ on each save and prunes older copies beyond this many |
prune_parts |
when TRUE (default), deletes the .parts/ resume cache after the canonical is written and verified; rebuilt from the canonical next run |
log_file |
path to a per-batch .rds log of what was fetched, when, and how many rows |
quiet |
suppresses headers, progress bars, and the info alerts |
polis_api_key |
the key; defaults to Sys.getenv("POLIS_API_KEY")
|
From download to clean
get_polis_data() writes the downloaded tables to disk; the cleaners take them from there. A typical flow pulls a table, then recovers any missing administrative geography from the EPID:
polished::get_polis_data(tables = "case")
cache <- tools::R_user_dir("polished", which = "cache")
cases <- readRDS(file.path(cache, "raw_afp.rds")) # `case` is saved as raw_afp
cleaned <- polished::impute_geo_from_epid(cases)
cleaned$qa # what was filled, and what was left unresolvedSee the Recovering geography from EPIDs vignette for that second step in detail, the End-to-end pipeline article to go from a folder of raw_* files straight to polished_* outputs, and ?get_polis_data for the complete argument reference.
