Skip to main content

Guides

From a website traffic API to a CSV you can actually compare

Run a complete Python script with real Vaneform API captures. Export monthly website traffic, preserve missing data, compare matching months and handle API errors.

Vaneform · Reviewed

A useful competitor CSV needs more than a domain and a traffic number. It needs the month that number describes, enough history to calculate change, and a visible result when a request has no usable data. This tutorial includes a complete standard-library Python script, three genuine API capture extracts and the resulting nine-row CSV. You can reproduce the published example before connecting your own key.

Reproduce the saved example without an account

Download the Python file and JSON extract into the same folder. Use Python 3.10 or later; there is no package installation step. The extract was recorded on 9 October 2026 in UTC+8 from successful /api/v1/domains/{domain}/scale requests for umami.is, simpleanalytics.com and screenshotone.com. It contains only the fields used here, kept unchanged from the responses. Original request and response credentials are excluded.

Run the command below from that folder. It writes nine rows: three domains, each with June, July and August 2026. The first available month has an empty previous_visits and mom_change_pct because May is absent from this capture. The script refuses to overwrite an existing output file, so use another filename when comparing a later run. The supplied CSV is available separately if you first want to inspect the result.

python3 traffic-to-csv.py \
  --snapshot traffic-api-snapshot-2026-10-09.json \
  --output traffic-replay.csv

# Wrote 9 rows; 0 unavailable/error rows.

The API response has two dates with different jobs

In the Simple Analytics response, scale.month is 2026-08, scale.as_of is 2026-08-01, estimate is site and confidence is medium. Its monthly_visits array contains 79,593, 113,086 and 107,563 visits for June through August. The displayed excerpt below is real captured data. The August first-day as_of value is supplied by the resource; it is preserved as provenance. The month field and each series point identify the measurement period.

The capture’s retrieved_at is a separate UTC timestamp for the HTTP read. It can fall on the previous UTC date relative to the UTC+8 editorial date. Keep all three concepts: measurement month, resource as_of and retrieval time. A request made in October can legitimately return August data, because this endpoint reads a stored resource with refresh_policy=cache_only. A successful request does not advance the data month.

The CSV repeats the resource’s confidence and as_of alongside the historical points so their provenance travels with the export. The response does not provide a separate confidence value for every historical month. Treat those columns as resource metadata, and keep the JSON if a later audit needs to recover the original shape. This export deliberately covers monthly site visits; it does not merge search estimates or channel-share fields into the series.

For a larger integration, keep the source response as the audit record and treat the CSV as a derived view. A field omitted from a successful response should stay absent or explicitly unavailable in your own model. The exporter checks that a usable site series includes its provenance and valid nonnegative numbers; it rejects duplicate months instead of choosing one arbitrarily. If the API contract changes, inspect the original payload and update the transformation deliberately.

{
  "month": "2026-08",
  "as_of": "2026-08-01",
  "confidence": "medium",
  "estimate": "site",
  "visits": 107563,
  "monthly_visits": [
    { "month": "2026-06", "visits": 79593 },
    { "month": "2026-07", "visits": 113086 },
    { "month": "2026-08", "visits": 107563 }
  ]
}

Connect your key when you want a fresh read of the stored resource

Open your Vaneform account and use its existing API key. Ordinary domain reads use the monthly API point allowance; a new usable domain read can consume a point. The documentation describes the shared 24-hour account-domain receipt and current plan limits. This example requests the scale resource only. It does not request Pro search expansion or trigger a bulk acquisition job.

In bash or zsh on macOS or Linux, the first command below waits for you to paste the key without echoing it. Press Enter, export the variable, then run the script. Keep the key out of the Python file, CSV and shared terminal transcripts. On another shell, set VANEFORM_API_KEY using that shell’s environment-variable mechanism and run the same Python arguments. A live response may contain a newer month or revised historical estimates than this saved example.

read -r -s VANEFORM_API_KEY
export VANEFORM_API_KEY
python3 traffic-to-csv.py umami.is simpleanalytics.com \
  --month 2026-08 --output traffic-live-2026-08.csv
unset VANEFORM_API_KEY

Compare one month without silently changing the sample

Add --month 2026-08 to the replay command to produce three August rows. For each domain, the script looks up the exact month in monthly_visits. It also looks up the immediately preceding calendar month, even if the array arrived unsorted. It does not treat the previous available observation as July when a month is missing. Without --month, it exports all recorded months, sorted within each domain.

The calculation is (current / previous − 1) × 100. For Simple Analytics, (107563 / 113086 − 1) × 100 rounds to −4.88%. mom_change_pct therefore stores −4.88, not the spreadsheet percentage fraction −0.0488. Import it as a number with a percent label in the column heading, or divide by 100 before applying percentage formatting. Otherwise a spreadsheet may display −488%, even though the export is correct.

A recorded zero remains zero. A missing baseline, or a zero baseline that makes division undefined, leaves growth blank. Filter row_status=ok and verify a common month before sorting by visits. The August result puts Umami first by estimated activity, followed by ScreenshotOne and Simple Analytics. That ordering is useful as a scale observation; choose comparable products before treating it as a competitive ranking.

Selected columns from the August replay; estimated visits, medium resource confidence.
domainmonthvisitsmom_change_pct
umami.is2026-0848183321.35
simpleanalytics.com2026-08107563-4.88
screenshotone.com2026-0812040111.34

Try an unavailable month before automating the export

Run the replay again with --month 2026-05 and a new output filename. The supplied snapshot has no May point, so the expected result is three month_unavailable rows with blank visits and exit status 1. Open that CSV next to the August result. The domains remain visible, while the date and reason tell you why they cannot enter a May total. This quick check is more useful than trusting a file simply because it has the expected number of rows.

When importing the file into a spreadsheet, set month to text and visits to a numeric column. Retain blank cells during import; a blanket “replace missing with zero” rule would erase the distinction this script preserves. The UTF-8 BOM helps common spreadsheet applications identify the file encoding. If you need JSON again later, use the saved extract rather than trying to reconstruct absent fields from the flatter CSV.

A repeatable job can retain one folder per retrieval date and compare coverage before comparing growth. If ten expected domains become eight usable rows, investigate the two failure states before reporting an aggregate decline. Save the script version alongside your job configuration, and run the offline example after changing the exporter. That gives you a known result to compare with without spending API points.

python3 traffic-to-csv.py \
  --snapshot traffic-api-snapshot-2026-10-09.json \
  --month 2026-05 --output traffic-missing-may.csv

# Wrote 3 rows; 3 unavailable/error rows.
# Exit status: 1

The failure rows are part of the output

A month missing from the series produces month_unavailable with an empty visit field. A missing or errored scale resource produces scale_unavailable. A search estimate produces unsupported_estimate because this file is for whole-site visits. A malformed history or missing provenance produces invalid_response. Keeping these rows prevents a failed candidate from quietly disappearing from the comparison and changing its denominator.

Authentication errors, credit exhaustion and other unsuccessful HTTP responses produce request_error with an error_code. The script makes at most three attempts for server errors 500/502/503/504 and lookup_rate_exceeded, with short bounded delays. It leaves credential and credit problems for you to resolve and rerun. Network failures remain recorded without an automatic retry. Redirects are refused so the bearer key stays attached only to the fixed Vaneform API origin.

Exit status 0 means every selected row is usable, 1 means the CSV includes unavailable or error rows, and 2 means the export could not be completed because of input, setup or output problems. In a scheduled job, check that exit code before sending the file to a dashboard. Save dated exports, inspect coverage changes, and keep the original month labels when appending a new snapshot. You then have a repeatable comparison rather than an unexplained column of numbers.

Research a website or a search opportunity · /guides/website-traffic-api-python-csv