# Pulls

A pull fetches an HTTPS endpoint on a schedule and stores what comes back in a source. It is the inbound counterpart to a drain: instead of you sending records to Tailglow, Tailglow goes and gets them. Use it for anything that already exposes its data over HTTP, such as a status endpoint, a JSON API, or a Prometheus metrics endpoint.

## Set up a pull

1. Open project **Settings**, then **Pulls**.
2. Click the plus button and pick the source the records land in.
3. Give the pull a name and the endpoint URL. The URL must be HTTPS and resolve to a public address.
4. Choose the method. **GET** suits most endpoints; **POST** is for endpoints that answer queries, such as GraphQL, and lets you supply a request body.
5. Add any headers the endpoint needs to authenticate you.
6. Choose a schedule.
7. Create the pull. It starts fetching straight away.

There is no verification handshake. A pull only reads from an endpoint you name, so there is nothing to prove to the far side.

## Headers

Header values are write-only. Tailglow encrypts them and never returns them, so the drawer lists the names you configured and shows the values as dots. To change one, add it again with the same name and the new value replaces it. To remove one, delete the row and save.

A pull sends at most 20 headers. Headers Tailglow sets itself, such as `Content-Type`, cannot be overridden.

## Schedule

Choose how often the endpoint is fetched: every minute, every 5 or 15 minutes, hourly, daily, or a custom cron expression. Schedules are evaluated in UTC and new pulls default to every minute, which is also the shortest supported gap.

Each run collects whatever the endpoint returns at that moment. A run that is missed, because the endpoint was slow or the schedule was paused, is not replayed later: the sample would carry the time it was collected rather than the time it was meant for, which would be worse than the gap.

## What gets stored

By default the whole response body becomes one record.

If the endpoint wraps its rows in a key, set **Records path** to that key and each element becomes its own record. `data` reads the top-level `data` array, and `data.result` walks two levels down. This is the difference between one record an hour containing five hundred rows, and five hundred records you can query.

Records land in a collection inside the source. Leave **Collection** empty and Tailglow names one after the pull.

## Watching an endpoint's availability

A pull tells you nothing about the runs that failed, because a failed run has nothing to store. It also stops after 72 hours of continuous failure: a collector that has been failing for three days has nothing left to collect.

That makes a pull the wrong tool for uptime. Create a [check](/guides/checks) instead. A check needs nothing but a URL, records one row per run whether the endpoint answered or not, and never stops itself because the endpoint is down, which is exactly the window you want a record of. If you want both the data and the availability, create both against the same URL.

## Prometheus endpoints

Pulls understand both Prometheus formats, and detect which one they received rather than asking you to declare it.

**Metrics endpoints.** A response in Prometheus text exposition format is parsed into one record per sample, with the metric name, labels, value and timestamp broken out as fields. You do not need a records path.

**Service discovery.** If the endpoint returns an HTTP service discovery document, the pull treats it as a list of other endpoints and fetches each of them. One scheduled run becomes one request to the discovery endpoint plus one request per target it names, and every sample is stored with the labels the discovery document attached to its target.

Two limits apply to discovery. A pull follows at most 1,000 targets, and the discovery document itself must be under 8 MB. Both refuse rather than truncate: a pull that quietly followed the first 1,000 of 4,000 targets would leave you with a metric that looks complete and is not. If you hit either, narrow what the discovery endpoint returns.

Your headers are sent to the discovery endpoint and to targets on the same origin. Targets on a different origin are fetched without them, so a token meant for one host is never handed to another.

## Testing and activity

**Test pull** on the Activity tab fetches the endpoint once, right now, and reports what happened: how many records came back, how long it took, and where it failed if it did. It does not store anything, so you can use it freely while you get the configuration right. A test that meets a discovery document follows at most five targets, enough to prove the shape without waiting for the full fan-out.

When a run fails, the pull records which step it failed at:

| Step       | Meaning                                                                 |
| ---------- | ----------------------------------------------------------------------- |
| `request`  | The endpoint could not be reached, refused the connection, or timed out |
| `response` | It answered with an error status, or a body too large to accept         |
| `parse`    | The body arrived but could not be read as JSON or Prometheus text       |
| `ingest`   | The records were read but could not be stored                           |

The Activity tab also shows the last 24 hours of volume, the most recent error, and how many runs have failed in a row.

## When something goes wrong

A failing pull keeps retrying on its schedule. Tailglow disables it after 72 hours of continuous failure and emails the team owners. Fix the endpoint and click **Resume**.

Pausing is yours and disabling is ours. A pull you paused stays paused until you resume it, and Tailglow never restarts it on your behalf. A pull Tailglow disabled says so, and resuming it is a deliberate act once the cause is fixed.

## Limits

Each request times out after 20 seconds. A response must fit the ingest payload limit, and a POST request body is capped at 4 KB. A pull follows at most 1,000 discovery targets, and a discovery document must be under 8 MB.

## Security

Endpoints must be public HTTPS addresses. Tailglow re-validates the address on every single fetch, not just when you save the pull, and rejects private, loopback and link-local targets. That re-validation is what stops a hostname that resolved publicly at save time from being repointed at an internal address later.
