Smalk Docs
  • Publisher
  • API Reference
  • Smalk MCP
  • Advertiser
Smalk
  • Website
  • Dashboard
  • Service status
Developers
  • REST API reference
  • OpenAPI schema
  • Support
Legal
  • Privacy policy
  • Terms

© 2026 Smalk. All rights reserved.

Information
Tracking
    Cloudflare Logpush ingestionpostCloudFront standard logs ingestion (Kinesis Firehose)postCloudFront S3 standard-log file ingestion (daily push)postClient-side trackingpostServer-side visit trackingpostBatched server-side visit trackingpost
GEA - Ad Content
Ad Placement Inventory
Workspace
Health
IndexNow
Reporting
Datasets
public
Schemas
powered by Zudoku
Smalk Public API
Smalk Public API

Tracking

Server-side and client-side tracking for AI agents


Cloudflare Logpush ingestion

POST
https://api.smalk.ai
/api/v1/tracking/cloudflare/logpush/

Ingest Cloudflare Logpush batches from the http_requests dataset.

Cloudflare Logpush delivers logs in small batches, potentially more than once per minute (Cloudflare Logpush docs).

Authentication

This endpoint uses the standard Smalk Project API key:

  • Header: Authorization: Api-Key <API_KEY>

Cloudflare’s HTTP destination configuration commonly injects headers via destination URL query parameters using the header_ convention (see destination configuration docs: Cloudflare destination_conf).

Example destination URL: https://api.smalk.ai/api/v1/tracking/cloudflare/logpush?header_Authorization=Api-Key%20<API_KEY>

Supported formats

  • NDJSON: one JSON object per line
  • JSON array: [{...}, {...}]
  • Optional gzip body (when Content-Encoding: gzip)

Recommended fields

At minimum, configure these fields in your Logpush job:

  • ClientRequestHost
  • ClientRequestMethod
  • ClientRequestReferer
  • ClientRequestURI
  • ClientRequestUserAgent
  • EdgeStartTimestamp (prefer RFC3339)

Do not add ClientIP: IP addresses are personal data under the GDPR and are neither processed nor stored. A record that still carries one has it filtered out at ingestion.

Cloudflare Logpush ingestion › Request Body

No data returned

Cloudflare Logpush ingestion › Responses

Batch accepted for processing

received
​integer · required

Total records received in the payload

queued
​integer · required

Records queued for processing (after validation)

POST/api/v1/tracking/cloudflare/logpush/
curl https://api.smalk.ai/api/v1/tracking/cloudflare/logpush \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{"ClientRequestHost":"example.com","ClientRequestMethod":"GET","ClientRequestURI":"/","ClientRequestReferer":"https://chat.openai.com/","ClientRequestUserAgent":"Mozilla/5.0 (compatible; ChatGPT-User/1.0)","EdgeStartTimestamp":"2025-12-18T12:34:56Z"} {"ClientRequestHost":"example.com","ClientRequestMethod":"GET","ClientRequestURI":"/pricing","ClientRequestReferer":"https://www.perplexity.ai/","ClientRequestUserAgent":"Mozilla/5.0 (compatible; PerplexityBot/1.0)","EdgeStartTimestamp":"2025-12-18T12:35:01Z"}'
Example Request Body
{"ClientRequestHost":"example.com","ClientRequestMethod":"GET","ClientRequestURI":"/","ClientRequestReferer":"https://chat.openai.com/","ClientRequestUserAgent":"Mozilla/5.0 (compatible; ChatGPT-User/1.0)","EdgeStartTimestamp":"2025-12-18T12:34:56Z"} {"ClientRequestHost":"example.com","ClientRequestMethod":"GET","ClientRequestURI":"/pricing","ClientRequestReferer":"https://www.perplexity.ai/","ClientRequestUserAgent":"Mozilla/5.0 (compatible; PerplexityBot/1.0)","EdgeStartTimestamp":"2025-12-18T12:35:01Z"}
json
application/json
Example Responses
{ "received": 100, "queued": 95 }
json
application/json

CloudFront standard logs ingestion (Kinesis Firehose)

POST
https://api.smalk.ai
/api/v1/tracking/cloudfront/firehose/

Ingest Amazon CloudFront standard logs (v2) delivered through a Kinesis Data Firehose HTTP-endpoint destination — no Kinesis Data Stream required.

CloudFront has no native "Logpush". The managed, no-code path is: CloudFront standard logging (v2) → Kinesis Data Firehose (HTTP endpoint destination, Output format = JSON) → this endpoint.

Authentication

The Smalk Project API key is carried in the Firehose access-key header:

  • Header: X-Amz-Firehose-Access-Key: <API_KEY>

(Firehose populates this from the destination's configured access key. If you use Secrets Manager, store it as {"api_key": "<API_KEY>"}.)

Request / response envelope

Firehose POSTs { "requestId", "timestamp", "records": [ {"data": "<base64>"} ] } (optionally gzip). Each base64 data blob decodes to one or more newline-separated CloudFront JSON log records. This endpoint always answers 200 with the Firehose ack { "requestId", "timestamp" } once auth passes — any non-2xx makes Firehose retry the whole batch.

CloudFront fields to select (standard logging v2, Output format = JSON)

Select at least these log fields on the delivery: date, time, cs-method, x-host-header, cs-uri-stem, cs-uri-query, cs(Referer), cs(User-Agent), sc-status.

Leave c-ip out: IP addresses are personal data under the GDPR and are neither processed nor stored. A record that still carries one has it filtered out at ingestion.

Field names are matched by name (no fixed order). W3C output (with a #Fields: header) is also accepted.

Once-a-day, no Firehose?

If you'd rather push logs on your own schedule without Firehose, send CloudFront S3 standard-log files to POST /api/v1/tracking/cloudfront/logs (Authorization: Api-Key).

CloudFront standard logs ingestion (Kinesis Firehose) › Request Body

No data returned

CloudFront standard logs ingestion (Kinesis Firehose) › Responses

Firehose ack (records accepted for processing)

requestId
​string · required

Echoed Firehose request id

timestamp
​integer · required

Server epoch milliseconds

POST/api/v1/tracking/cloudfront/firehose/
curl https://api.smalk.ai/api/v1/tracking/cloudfront/firehose \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{}'
Example Request Body
{}
json
Example Responses
{ "requestId": "requestId", "timestamp": 0 }
json
application/json

CloudFront S3 standard-log file ingestion (daily push)

POST
https://api.smalk.ai
/api/v1/tracking/cloudfront/logs/

Ingest Amazon CloudFront standard log files for publishers who want to push logs on their own schedule (e.g. once a day) with no Firehose and no Kinesis Data Stream.

Enable CloudFront standard logging → Amazon S3, then run a scheduled job (Lambda / cron) that reads the day's gzipped log objects and forwards the bytes to this endpoint.

Authentication

  • Header: Authorization: Api-Key <API_KEY>

Body

One or more CloudFront S3 standard-log objects:

  • W3C text (the native S3 format) with a #Fields: header line — column order is read from that header, gzip accepted (Content-Encoding: gzip or sniffed).
  • or NDJSON / JSON-array of CloudFront JSON records.

Static assets (.js/.css/fonts/images/media) are dropped at ingest. Max 20MB per request — POST per S3 object (or batch several) to stay under the limit.

CloudFront S3 standard-log file ingestion (daily push) › Request Body

No data returned

CloudFront S3 standard-log file ingestion (daily push) › Responses

Logs accepted for processing

No data returned
POST/api/v1/tracking/cloudfront/logs/
curl https://api.smalk.ai/api/v1/tracking/cloudfront/logs \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{}'
Example Request Body
{}
json
Example Responses
No example specified for this content type

Client-side tracking

POST
https://api.smalk.ai
/api/v1/tracking/track/

Record a page visit from the browser. This endpoint is typically called by the Smalk JavaScript tracker (tracker.js) to capture client-side visits.

Authentication: None required (uses workspace_key in request body).

Important Note

This endpoint only tracks visitors that execute JavaScript. AI agents (ChatGPT, Perplexity, etc.) don't execute JavaScript, so you should also implement server-side tracking via POST /api/v1/tracking/visit/ for complete AI agent coverage.

Use Cases

  • Browser-based tracking with full user context
  • Capturing screen size, navigator info, and session data
  • Integration with single-page applications (SPAs)
  • Human visitor analytics

Client-side tracking › Request Body

project_key
​string · uuid · required
​object · required
path
​string · maxLength: 2048 · required

Client-side tracking › Responses

Track event accepted for processing

No data returned
POST/api/v1/tracking/track/
curl https://api.smalk.ai/api/v1/tracking/track \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{ "project_key": "550e8400-e29b-41d4-a716-446655440000", "path": "/products/widget", "browser_session_id": "660e8400-e29b-41d4-a716-446655440001", "tracker_request_headers": { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" }, "browser_screen": { "width": 1920, "height": 1080 }, "browser_navigator": { "language": "en-US", "platform": "Win32" } }'
Example Request Body
{ "project_key": "550e8400-e29b-41d4-a716-446655440000", "path": "/products/widget", "browser_session_id": "660e8400-e29b-41d4-a716-446655440001", "tracker_request_headers": { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" }, "browser_screen": { "width": 1920, "height": 1080 }, "browser_navigator": { "language": "en-US", "platform": "Win32" } }
json
Example Responses
No example specified for this content type

Server-side visit tracking

POST
https://api.smalk.ai
/api/v1/tracking/visit/

Record a page visit from your server. This endpoint is essential for tracking AI agents (ChatGPT, Perplexity, Claude, Gemini, etc.) as they don't execute JavaScript.

Authentication: Requires API Key in the Authorization header.

Why Server-Side Tracking is Critical

AI agents and LLM scrapers don't load frontend JavaScript, so the backend integration is essential to track them. Both frontend and backend integrations are recommended for complete tracking coverage.

Request Headers Guidelines

The request_headers field should contain only the headers we need for AI agent detection and analytics. Please filter your headers before sending.

✅ Required Headers

HeaderDescription
User-AgentRequired. The visitor's User-Agent string. This is how we identify AI agents.
RefererRequired. Referrer URL to track traffic sources (e.g., from ChatGPT, Perplexity).

🔌 CMS Plugin Headers

If you are building a CMS plugin, include these HTTP headers on each request to enable automatic publisher installation tracking (updated once per day):

HTTP HeaderDescriptionExample
X-Smalk-CMSRecommended. CMS type and version. Format: {cms}/{version}.wordpress/6.8, drupal/10.3.1
X-Smalk-Plugin-VersionRecommended. Your Smalk plugin version.1.0.10

⛔ Do NOT Include These Headers

Important: For security reasons, do not forward sensitive headers to our API. Filter out these headers before sending:

  • Authorization - Your internal auth tokens
  • Cookie - Session cookies
  • X-Auth-Token, X-API-Key - Any authentication tokens
  • JWT, Bearer tokens - Any JWT or bearer authentication
  • X-CSRF-Token - CSRF tokens
  • Set-Cookie - Cookie setting headers
  • Any custom authentication or session headers

🔒 GDPR Compliance Guarantee: If you accidentally send any headers outside our required/recommended list, we guarantee they will not be analyzed, processed, or stored in any way. We only extract and store data from the headers listed above to maintain full GDPR compliance.

Example of header filtering (Python):

Code
# IP-bearing headers (x-real-ip, x-forwarded-for) are dropped server-side for # GDPR — no need to forward them. ALLOWED_HEADERS = {'user-agent', 'referer'} filtered_headers = {k: v for k, v in request.headers.items() if k.lower() in ALLOWED_HEADERS}

Use Cases

  • AI Agent Tracking - ChatGPT, Perplexity, Claude, Gemini, and other LLM crawlers
  • Server-side analytics - For pages that don't load JavaScript
  • Framework Integration - WordPress, Next.js, Django, Rails, or any server framework

Best Practices

1. Feature Flag Integration

Add a feature flag around the server call. This allows you to quickly disable tracking if needed, providing flexibility and control over your analytics.

2. Timeout Configuration

We guarantee a backend API response time of less than 100ms. We recommend setting a short timeout of around 150ms to ensure your pages don't experience any noticeable slowdown.

3. Use Async/Non-Blocking Calls

Trigger the tracking request asynchronously (fire-and-forget) so it never blocks your main response path. Use background workers, edge workers, or non-awaited promises with short timeouts.

Response Codes

  • 202 - Visit accepted for processing
  • 204 - Path is excluded from tracking (expected, no action needed)
  • 400 - Validation error in request data
  • 401 - Invalid or missing API Key

Server-side visit tracking › Request Body

request_path
​string · maxLength: 2048 · required

The URL path being visited (e.g., '/blog/my-article')

request_method
​string · enum · required

HTTP method of the request (GET, POST, etc.)

  • GET - GET
  • POST - POST
  • PUT - PUT
  • HEAD - HEAD
  • DELETE - DELETE
  • PATCH - PATCH
  • OPTIONS - OPTIONS
  • TRACE - TRACE
Enum values:
GET
POST
PUT
HEAD
DELETE
PATCH
OPTIONS
TRACE
​object · required

Filtered request headers. Only include: User-Agent (required), Referer (recommended). IP-bearing headers (X-Real-IP, X-Forwarded-For) are dropped server-side for GDPR and need not be sent. Do NOT include sensitive headers like Authorization, Cookie, or tokens.

timestamp
​

Server-side visit tracking › Responses

Visit accepted for processing

No data returned
POST/api/v1/tracking/visit/
curl https://api.smalk.ai/api/v1/tracking/visit \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{ "request_path": "/blog/my-article", "request_method": "GET", "request_headers": { "User-Agent": "Mozilla/5.0 (compatible; ChatGPT-User/1.0)", "Referer": "https://chat.openai.com/" } }'
Example Request Body
{ "request_path": "/blog/my-article", "request_method": "GET", "request_headers": { "User-Agent": "Mozilla/5.0 (compatible; ChatGPT-User/1.0)", "Referer": "https://chat.openai.com/" } }
json
Example with only the required and recommended headers
Example Responses
No example specified for this content type

Batched server-side visit tracking

POST
https://api.smalk.ai
/api/v1/tracking/visit/batch/

Accept up to 500 visits in a single request. Behaviour for each visit is identical to the single POST /api/v1/tracking/visit/ endpoint — bot detection, path exclusion, and analytics aggregation all apply per visit — amortized over a batch so high-traffic sites can flush buffered visits with a single HTTP call.

Designed for CMS plugins that buffer visits locally and flush them on a periodic interval.

Authentication: Requires API Key in the Authorization header. Format: Authorization: Api-Key <API_KEY>

Headers

HeaderPurpose
X-Smalk-CMSOptional, e.g. wordpress/6.9.4 — recorded with the batch so dashboards know which CMS sent the visit
X-Smalk-Plugin-VersionOptional — recorded with the batch
X-Smalk-Probe: 1Optional debug probe. Returns {probe: true} without recording any analytics — use this to verify connectivity from support tooling without polluting your real analytics

Per-visit exclusion: each visit's request_path is checked against the project's exclusion rules; matching visits are silently dropped and not counted in queued.

Per-visit timestamp: each visit can carry an optional timestamp field (ISO 8601 like 2026-05-22T08:14:33Z, or a Unix-seconds integer). Set it to the moment the client observed the visit and Smalk will store the event at that time rather than when the batch reached the server. Useful when the client buffers visits locally and flushes them later. When omitted, the server falls back to the batch arrival time.

Body schema

Code
{ "visits": [ { "request_path": "/page", "request_method": "GET", "request_headers": {"user-agent": "GPTBot/1.0", "referer": "..."}, "timestamp": "2026-05-22T08:14:33Z" } ] }

Response: 202 Accepted with an empty body. The endpoint is fire-and-forget; clients should not parse the response. Drop the connection as soon as the 202 lands. Error responses (4xx) carry a JSON body with a detail field.

Batched server-side visit tracking › Request Body

No data returned

Batched server-side visit tracking › Responses

Batch accepted (empty body — fire-and-forget).

No data returned
POST/api/v1/tracking/visit/batch/
curl https://api.smalk.ai/api/v1/tracking/visit/batch \ --request POST \ --header 'Content-Type: application/json' \ --header 'Authorization: <api-key>' \ --data '{}'
Example Request Body
{}
json
Example Responses
No example specified for this content type

GEA - Ad Content
JSON