Tracking
Server-side and client-side tracking for AI agents
Cloudflare Logpush ingestion
Ingest Cloudflare Logpush batches from the http_requests dataset.
Cloudflare Logpush delivers logs in small batches, potentially more than once per minute (Cloudflare Logpush docs).
Authentication
This endpoint uses the standard Smalk Project API key:
- Header:
Authorization: Api-Key <API_KEY>
Cloudflare’s HTTP destination configuration commonly injects headers via destination URL
query parameters using the header_ convention (see destination configuration docs:
Cloudflare destination_conf).
Example destination URL:
https://api.smalk.ai/api/v1/tracking/cloudflare/logpush?header_Authorization=Api-Key%20<API_KEY>
Supported formats
- NDJSON: one JSON object per line
- JSON array:
[{...}, {...}] - Optional gzip body (when
Content-Encoding: gzip)
Recommended fields
At minimum, configure these fields in your Logpush job:
ClientRequestHostClientRequestMethodClientRequestRefererClientRequestURIClientRequestUserAgentEdgeStartTimestamp(prefer RFC3339)
Do not add ClientIP: IP addresses are personal data under the GDPR and are neither
processed nor stored. A record that still carries one has it filtered out at ingestion.
Cloudflare Logpush ingestion › Responses
Batch accepted for processing
receivedTotal records received in the payload
queuedRecords queued for processing (after validation)
CloudFront standard logs ingestion (Kinesis Firehose)
Ingest Amazon CloudFront standard logs (v2) delivered through a Kinesis Data Firehose HTTP-endpoint destination — no Kinesis Data Stream required.
CloudFront has no native "Logpush". The managed, no-code path is: CloudFront standard logging (v2) → Kinesis Data Firehose (HTTP endpoint destination, Output format = JSON) → this endpoint.
Authentication
The Smalk Project API key is carried in the Firehose access-key header:
- Header:
X-Amz-Firehose-Access-Key: <API_KEY>
(Firehose populates this from the destination's configured access key. If you
use Secrets Manager, store it as {"api_key": "<API_KEY>"}.)
Request / response envelope
Firehose POSTs { "requestId", "timestamp", "records": [ {"data": "<base64>"} ] }
(optionally gzip). Each base64 data blob decodes to one or more
newline-separated CloudFront JSON log records. This endpoint always answers
200 with the Firehose ack { "requestId", "timestamp" } once auth passes —
any non-2xx makes Firehose retry the whole batch.
CloudFront fields to select (standard logging v2, Output format = JSON)
Select at least these log fields on the delivery:
date, time, cs-method, x-host-header, cs-uri-stem,
cs-uri-query, cs(Referer), cs(User-Agent), sc-status.
Leave c-ip out: IP addresses are personal data under the GDPR and are neither
processed nor stored. A record that still carries one has it filtered out at ingestion.
Field names are matched by name (no fixed order). W3C output (with a #Fields:
header) is also accepted.
Once-a-day, no Firehose?
If you'd rather push logs on your own schedule without Firehose, send CloudFront
S3 standard-log files to POST /api/v1/tracking/cloudfront/logs
(Authorization: Api-Key).
CloudFront standard logs ingestion (Kinesis Firehose) › Responses
Firehose ack (records accepted for processing)
requestIdEchoed Firehose request id
timestampServer epoch milliseconds
CloudFront S3 standard-log file ingestion (daily push)
Ingest Amazon CloudFront standard log files for publishers who want to push logs on their own schedule (e.g. once a day) with no Firehose and no Kinesis Data Stream.
Enable CloudFront standard logging → Amazon S3, then run a scheduled job (Lambda / cron) that reads the day's gzipped log objects and forwards the bytes to this endpoint.
Authentication
- Header:
Authorization: Api-Key <API_KEY>
Body
One or more CloudFront S3 standard-log objects:
- W3C text (the native S3 format) with a
#Fields:header line — column order is read from that header, gzip accepted (Content-Encoding: gzipor sniffed). - or NDJSON / JSON-array of CloudFront JSON records.
Static assets (.js/.css/fonts/images/media) are dropped at ingest. Max 20MB
per request — POST per S3 object (or batch several) to stay under the limit.
CloudFront S3 standard-log file ingestion (daily push) › Responses
Logs accepted for processing
Client-side tracking
Record a page visit from the browser. This endpoint is typically called by
the Smalk JavaScript tracker (tracker.js) to capture client-side visits.
Authentication: None required (uses workspace_key in request body).
Important Note
This endpoint only tracks visitors that execute JavaScript. AI agents
(ChatGPT, Perplexity, etc.) don't execute JavaScript, so you should also
implement server-side tracking via POST /api/v1/tracking/visit/ for
complete AI agent coverage.
Use Cases
- Browser-based tracking with full user context
- Capturing screen size, navigator info, and session data
- Integration with single-page applications (SPAs)
- Human visitor analytics
Client-side tracking › Request Body
project_keypathClient-side tracking › Responses
Track event accepted for processing
Server-side visit tracking
Record a page visit from your server. This endpoint is essential for tracking AI agents (ChatGPT, Perplexity, Claude, Gemini, etc.) as they don't execute JavaScript.
Authentication: Requires API Key in the Authorization header.
Why Server-Side Tracking is Critical
AI agents and LLM scrapers don't load frontend JavaScript, so the backend integration is essential to track them. Both frontend and backend integrations are recommended for complete tracking coverage.
Request Headers Guidelines
The request_headers field should contain only the headers we need for AI agent
detection and analytics. Please filter your headers before sending.
✅ Required Headers
| Header | Description |
|---|---|
User-Agent | Required. The visitor's User-Agent string. This is how we identify AI agents. |
Referer | Required. Referrer URL to track traffic sources (e.g., from ChatGPT, Perplexity). |
🔌 CMS Plugin Headers
If you are building a CMS plugin, include these HTTP headers on each request to enable automatic publisher installation tracking (updated once per day):
| HTTP Header | Description | Example |
|---|---|---|
X-Smalk-CMS | Recommended. CMS type and version. Format: {cms}/{version}. | wordpress/6.8, drupal/10.3.1 |
X-Smalk-Plugin-Version | Recommended. Your Smalk plugin version. | 1.0.10 |
⛔ Do NOT Include These Headers
Important: For security reasons, do not forward sensitive headers to our API. Filter out these headers before sending:
Authorization- Your internal auth tokensCookie- Session cookiesX-Auth-Token,X-API-Key- Any authentication tokensJWT,Bearertokens - Any JWT or bearer authenticationX-CSRF-Token- CSRF tokensSet-Cookie- Cookie setting headers- Any custom authentication or session headers
🔒 GDPR Compliance Guarantee: If you accidentally send any headers outside our required/recommended list, we guarantee they will not be analyzed, processed, or stored in any way. We only extract and store data from the headers listed above to maintain full GDPR compliance.
Example of header filtering (Python):
Code
Use Cases
- AI Agent Tracking - ChatGPT, Perplexity, Claude, Gemini, and other LLM crawlers
- Server-side analytics - For pages that don't load JavaScript
- Framework Integration - WordPress, Next.js, Django, Rails, or any server framework
Best Practices
1. Feature Flag Integration
Add a feature flag around the server call. This allows you to quickly disable tracking if needed, providing flexibility and control over your analytics.
2. Timeout Configuration
We guarantee a backend API response time of less than 100ms. We recommend setting a short timeout of around 150ms to ensure your pages don't experience any noticeable slowdown.
3. Use Async/Non-Blocking Calls
Trigger the tracking request asynchronously (fire-and-forget) so it never blocks your main response path. Use background workers, edge workers, or non-awaited promises with short timeouts.
Response Codes
- 202 - Visit accepted for processing
- 204 - Path is excluded from tracking (expected, no action needed)
- 400 - Validation error in request data
- 401 - Invalid or missing API Key
Server-side visit tracking › Request Body
request_pathThe URL path being visited (e.g., '/blog/my-article')
request_methodHTTP method of the request (GET, POST, etc.)
GET- GETPOST- POSTPUT- PUTHEAD- HEADDELETE- DELETEPATCH- PATCHOPTIONS- OPTIONSTRACE- TRACE
Filtered request headers. Only include: User-Agent (required), Referer (recommended). IP-bearing headers (X-Real-IP, X-Forwarded-For) are dropped server-side for GDPR and need not be sent. Do NOT include sensitive headers like Authorization, Cookie, or tokens.
timestampServer-side visit tracking › Responses
Visit accepted for processing
Batched server-side visit tracking
Accept up to 500 visits in a single request. Behaviour for each visit is
identical to the single POST /api/v1/tracking/visit/ endpoint — bot
detection, path exclusion, and analytics aggregation all apply per visit —
amortized over a batch so high-traffic sites can flush buffered visits with
a single HTTP call.
Designed for CMS plugins that buffer visits locally and flush them on a periodic interval.
Authentication: Requires API Key in the Authorization header.
Format: Authorization: Api-Key <API_KEY>
Headers
| Header | Purpose |
|---|---|
X-Smalk-CMS | Optional, e.g. wordpress/6.9.4 — recorded with the batch so dashboards know which CMS sent the visit |
X-Smalk-Plugin-Version | Optional — recorded with the batch |
X-Smalk-Probe: 1 | Optional debug probe. Returns {probe: true} without recording any analytics — use this to verify connectivity from support tooling without polluting your real analytics |
Per-visit exclusion: each visit's request_path is checked against the
project's exclusion rules; matching visits are silently dropped and not
counted in queued.
Per-visit timestamp: each visit can carry an optional timestamp field
(ISO 8601 like 2026-05-22T08:14:33Z, or a Unix-seconds integer). Set it to
the moment the client observed the visit and Smalk will store the event at
that time rather than when the batch reached the server. Useful when the
client buffers visits locally and flushes them later. When omitted, the
server falls back to the batch arrival time.
Body schema
Code
Response: 202 Accepted with an empty body. The endpoint is fire-and-forget;
clients should not parse the response. Drop the connection as soon as the
202 lands. Error responses (4xx) carry a JSON body with a detail field.
Batched server-side visit tracking › Responses
Batch accepted (empty body — fire-and-forget).