# Pulls

## Pull Model

### Fields

- **`object`** `"pull"`

   Always `"pull"`.

- **`id`** `string`

   Unique identifier, prefixed with `pul_`.

- **`project_id`** `string`

   The project this pull belongs to.

- **`team_id`** `string`

   The team that owns the pull.

- **`name`** `string`

   Display name for the pull.

- **`source_id`** `string`

   The source fetched records are written to.

- **`source_name`** `string | null`

   Display name of that source.

- **`collection_slug`** `string | null`

   The collection within the source that records land in. Defaults to a slug of the pull name.

- **`endpoint_url`** `string`

   The HTTPS URL that is fetched on each run.

- **`method`** [`PullMethod`](/api/pulls#pull-method)

   The HTTP method used to fetch the endpoint.

- **`headers_configured`** `string[]`

   Names of the headers sent with every fetch. The values are stored encrypted and are never
   returned, so this shows what was configured without revealing the secrets.

- **`request_body`** `string | null`

   The request body sent when the method is `post`. Null for `get`.

- **`json_records_path`** `string | null`

   Dot path to the array of records inside a JSON response, such as `data` or `data.result`. Null
   when the response is stored exactly as it arrives.

- **`schedule_cron`** `string`

   Cron expression setting when the endpoint is fetched, evaluated in UTC.

- **`status`** [`PullStatus`](/api/pulls#pull-status)

   The current state of the pull.

- **`next_fetch_at`** [`ISODateString | null`](/api/pulls#iso-date-string)

   When the next fetch is due. Null when the pull is paused or disabled, because nothing is
   scheduled: a time here would promise a fetch that will not happen.

- **`missed_slot_count`** `number`

   How many scheduled slots have elapsed without a fetch, since the pull was created or last
   resumed. Almost always because the project had no healthy server at the time.
   
   Missed slots are counted rather than fetched late. A pull records whatever its endpoint serves
   at the moment of the request, so a delayed fetch would return current data stamped with a time
   it does not describe. A rising count next to a healthy `last_success_at` means the schedule is
   finer than the project's capacity has been able to keep up with.

- **`last_attempt_at`** [`ISODateString | null`](/api/pulls#iso-date-string)

   When the endpoint was last fetched, whether or not it succeeded.

- **`last_success_at`** [`ISODateString | null`](/api/pulls#iso-date-string)

   When records were last fetched and accepted.

- **`last_failure_at`** [`ISODateString | null`](/api/pulls#iso-date-string)

   When a fetch last failed.

- **`consecutive_failure_count`** `number`

   How many fetches have failed in a row. Resets to zero on the next success.

- **`last_error`** `string | null`

   The error from the most recent failed fetch.

- **`last_error_stage`** [`PullErrorStage | null`](/api/pulls#pull-error-stage)

   Which step of the most recent failed fetch went wrong.

- **`last_http_status`** `number | null`

   The HTTP status the endpoint returned on the most recent fetch.

- **`last_response_time_ms`** `number | null`

   How long the most recent fetch took, in milliseconds.

- **`last_record_count`** `number`

   How many records the most recent successful fetch produced.

- **`records_pulled_hourly`** `PullHourlyRecords[]`

   The last 24 hours of fetch volume, oldest first. Always 24 entries, so an hour with no fetches
   reads as a zero rather than being absent, and a pull that has never run still returns a full
   window of zeros. Sum the counts for "records pulled in the last 24 hours".

- **`created_at`** [`ISODateString`](/api/pulls#iso-date-string)

   When the pull was created.

- **`updated_at`** [`ISODateString`](/api/pulls#iso-date-string)

   When the pull was last changed.

### Referenced Types

#### ISODateString

`ISODateString`

An ISO 8601 date-time string returned at the JSON API boundary.

#### PullMethod

`"get" | "post"`

#### PullStatus

`"active" | "paused" | "disabled"`

#### PullErrorStage

`"request" | "response" | "parse" | "ingest"`

## Pull Test Result

### Fields

- **`ok`** `boolean`

   Whether the endpoint responded successfully and the response could be read as records.

- **`status`** `number`

   The HTTP status the endpoint returned. Zero when no response was received at all.

- **`status_text`** `string`

   The HTTP status text the endpoint returned.

- **`response_time_ms`** `number`

   How long the fetch took, in milliseconds.

- **`byte_length`** `number`

   How many bytes the endpoint returned.

- **`record_count`** `number`

   How many records the response would produce. Zero when the test failed.

- **`body_preview`** `string`

   The first 2 KB of the response body.

- **`error_stage`** [`PullErrorStage | null`](/api/pulls#pull-error-stage)

   Which step went wrong. Null when the test succeeded.

- **`error`** `string | null`

   What went wrong. Null when the test succeeded.

### Referenced Types

#### ISODateString

`ISODateString`

An ISO 8601 date-time string returned at the JSON API boundary.

#### PullErrorStage

`"request" | "response" | "parse" | "ingest"`

## List Pulls

### Endpoint

Retrieve a list of pulls.

```http
GET /v1/projects/:project_id/pulls
```

**Scope:** `pulls:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

### Query Parameters

- **`order_by`** `string`
  Field the results are ordered by. Optional. Defaults to `"created_at"`. Allowed values: `"name"`, `"created_at"`, `"last_success_at"`.

- **`status`** `string`
  Return only pulls in this state. Optional. Allowed values: `"active"`, `"paused"`, `"disabled"`.

- **`source_id`** `string`
  Return only pulls that write to this source. Optional.

- **`limit`** `number`
  Maximum number of items to return. Optional. Defaults to `25`. Minimum: `1`. Maximum: `200`.

- **`after`** `string`
  Cursor from `pagination.next_cursor` of a previous response. Returns the resources after that page. Optional.

- **`before`** `string`
  Cursor from `pagination.prev_cursor` of a previous response. Returns the resources before that page. Optional.

- **`sort`** `string`
  Sort direction for the result set. Optional. Defaults to `"desc"`. Allowed values: `"asc"`, `"desc"`.

### Response

```ts
{
  message: string;
  data: Pull[];
  status: 200;
  error: null;
  pagination: Pagination;
  endpoint: string;
}
```

### Comments

- `after` and `before` are mutually exclusive.

## Retrieve Pull

### Endpoint

Retrieve a single pull.

```http
GET /v1/projects/:project_id/pulls/:pull_id
```

**Scope:** `pulls:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Response

Pull retrieved

```ts
{
  message: string;
  data: Pull;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

## Create Pull

### Endpoint

Create a scheduled pull in a project.

```http
POST /v1/projects/:project_id/pulls
```

**Scope:** `pulls:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

### Request Body

- **`name`** `string` -- **Required**
  Name shown for this pull. Minimum length: `2`. Maximum length: `128`.

- **`source_id`** `string` -- **Required**
  Source that fetched records are written to. Minimum length: `1`.

- **`collection_slug`** `string | null`
  Collection within the source that records land in. Leave unset and it is derived from the pull name, so each pull gets its own stream. Optional.

- **`endpoint_url`** `string` -- **Required**
  HTTPS URL fetched on every run. Must resolve to a public address. Minimum length: `1`. Maximum length: `2048`.

- **`method`** `string`
  HTTP method used to fetch the endpoint. Optional. Defaults to `"get"`. Allowed values: `"get"`, `"post"`.

- **`headers`** `object | null`
  Headers sent with every fetch, for authenticating to the endpoint. Stored encrypted and never returned; responses list only the header names, as `headers_configured`. Optional.

- **`request_body`** `string | null`
  Body sent with every fetch. Only valid when `method` is `post`, for endpoints that answer queries over POST. Optional.

- **`json_records_path`** `string | null`
  Dot path to the array of records inside a JSON response, for example `data` or `data.result`. Leave unset to store the response exactly as it arrives. Optional.

- **`schedule_cron`** `string`
  Cron expression setting when the endpoint is fetched, evaluated in UTC. Defaults to every minute, which is also the shortest supported gap. Optional. Defaults to `"* * * * *"`. Minimum length: `1`. Maximum length: `128`.

### Response

Pull created

```ts
{
  message: string;
  data: Pull;
  status: 201;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Pull names are unique within a project.
- The endpoint must be reachable over HTTPS at a public address. Addresses inside private or reserved ranges are rejected, and the check is repeated on every fetch, not just at create.
- A pull is created `active` even if the project has no servers yet. It starts fetching as soon as a server exists; an empty fleet delays the first fetch rather than changing the status.
- An endpoint answering with a Prometheus HTTP Service Discovery document is followed rather than stored. Each discovered target is fetched, its metrics parsed into records, and the target's labels merged onto them so targets stay distinguishable. This is what makes an endpoint that hands out short-lived signed URLs work on a schedule: the discovery response is re-read on every run, so the credentials are always current.
- Discovered targets are fetched over HTTPS at public addresses only, checked with the same guard as the endpoint itself, so a discovery response cannot direct a pull at a private address.
- A pull follows at most 50 discovered targets and fails if the document lists more, rather than fetching some and silently omitting the rest. Individual targets are retried; if some still fail, the run succeeds with the records it did collect and reports how many targets did not.
- `request_body` is only accepted when `method` is `post`. Sending one with a `get` pull is rejected rather than ignored, so a half-configured pull fails at create time instead of quietly fetching the wrong thing.
- `headers` accepts at most 20 entries. A header whose value is `null` is ignored here, since on create there is nothing yet for it to remove. Names are matched case-insensitively, so two spellings of one name are stored as a single header.
- `schedule_cron` must be a valid cron expression. It is evaluated in UTC, and the finest granularity is one minute.
- `collection_slug` must be lowercase, start with a letter or digit, and may otherwise contain digits, letters, hyphens and underscores.
- `json_records_path` is for endpoints that wrap their rows in an envelope. An endpoint returning `{"data": [{...}, {...}]}` stores one record containing the whole response unless you set the path to `data`, which stores the two records instead. Nested keys are joined with dots (`data.result`). Object keys only: array indexes and wildcards are not supported.
- `json_records_path` applies only to JSON responses. JSONL, CSV, TSV and Prometheus responses already produce one record per row, so leave it unset for those.
- A fetch fails with a `parse` error when the path is missing from the response or does not resolve to a list or object, rather than falling back to storing the whole document. Use the test endpoint to check a path against a live response before saving.

## Update Pull

### Endpoint

Update a pull.

```http
POST /v1/projects/:project_id/pulls/:pull_id
```

**Scope:** `pulls:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Request Body

- **`name`** `string`
  Name shown for this pull. Optional. Minimum length: `2`. Maximum length: `128`.

- **`collection_slug`** `string | null`
  Collection within the source that records land in. Leave unset and it is derived from the pull name, so each pull gets its own stream. Optional.

- **`endpoint_url`** `string`
  HTTPS URL fetched on every run. Must resolve to a public address. Optional. Minimum length: `1`. Maximum length: `2048`.

- **`method`** `string`
  HTTP method used to fetch the endpoint. Optional. Allowed values: `"get"`, `"post"`.

- **`headers`** `object | null`
  Headers sent with every fetch, for authenticating to the endpoint. Stored encrypted and never returned; responses list only the header names, as `headers_configured`. Optional.

- **`request_body`** `string | null`
  Body sent with every fetch. Only valid when `method` is `post`, for endpoints that answer queries over POST. Optional.

- **`json_records_path`** `string | null`
  Dot path to the array of records inside a JSON response, for example `data` or `data.result`. Leave unset to store the response exactly as it arrives. Optional.

- **`schedule_cron`** `string`
  Cron expression setting when the endpoint is fetched, evaluated in UTC. Defaults to every minute, which is also the shortest supported gap. Optional. Minimum length: `1`. Maximum length: `128`.

### Response

Pull updated

```ts
{
  message: string;
  data: Pull;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Changes take effect on the next fetch; a fetch already in flight finishes under the old configuration.
- A changed `schedule_cron` takes effect from the next fetch onward.
- At least one field must be provided.
- Any field not listed here is rejected, including `source_id`, which cannot be changed after create. Create a new pull to write into a different source.
- Clearing `request_body` is required before changing `method` from `post` to `get`; a pull cannot keep a body it would never send.
- `headers` is a patch, not a replacement, because the values are write-only and never returned. A name mapped to a string adds or replaces that header, a name mapped to `null` removes it, and a name you do not mention keeps its stored value. Omit the field to leave every header alone, or send `null` in place of the object to remove all of them.
- `headers` names are matched case-insensitively, so patching `authorization` replaces a stored `Authorization` rather than adding a second header. The spelling you send is the one stored and sent.
- `headers` accepts at most 20 entries, counted after the patch is applied.
- `schedule_cron` must be a valid cron expression, evaluated in UTC.
- Send `json_records_path` as `null` to stop unwrapping and store responses whole again.

## Delete Pull

### Endpoint

Delete a pull.

```http
DELETE /v1/projects/:project_id/pulls/:pull_id
```

**Scope:** `pulls:delete`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Response

Pull deleted

```ts
{
  message: string;
  data: null;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Records the pull already wrote are not deleted. They belong to the source and stay queryable.

## Pause Pull

### Endpoint

Pause a pull so it stops fetching.

```http
POST /v1/projects/:project_id/pulls/:pull_id/pause
```

**Scope:** `pulls:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Response

Pull paused

```ts
{
  message: string;
  data: Pull;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Only an active pull can be paused. Pausing an already-paused pull succeeds and changes nothing.

## Resume Pull

### Endpoint

Resume a paused pull.

```http
POST /v1/projects/:project_id/pulls/:pull_id/resume
```

**Scope:** `pulls:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Response

Pull resumed

```ts
{
  message: string;
  data: Pull;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Resumes a pull you paused, and also re-enables one the system disabled after a long failure streak. Resuming clears the failure count, so it starts from a clean slate.

## Test Pull

### Endpoint

Fetch the endpoint once without storing anything.

```http
POST /v1/projects/:project_id/pulls/:pull_id/test
```

**Scope:** `pulls:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`pull_id`** `string` -- **Required**
  Unique identifier of the pull.

### Response

```ts
{
  message: string;
  data: PullTestResult;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Nothing is ingested and no health field or schedule is changed, so this is safe to run against a live pull.
- The response reports the HTTP status, how long the fetch took, how many records the configured format would produce, and the first 2 KB of the body.

