# Collections

## Collection Model

### Fields

- **`object`** `"collection"`

- **`id`** `string`

   Unique identifier, prefixed with `col_`.

- **`source_id`** `string`

- **`source_name`** `string | null`

- **`project_id`** `string`

- **`slug`** `string`

   Permanent identifier used in endpoints. Cannot be changed after creation.

- **`name`** `string`

- **`storage_gb`** `number`

   Data stored in this collection's files, in gigabytes of uncompressed data. Recalculated hourly.

- **`schema_count`** `number`

   How many schema versions this collection has, in every state.

- **`views_count`** `number`

   How many views read from this collection.

- **`drains_count`** `number`

   How many drains export from this collection.

- **`schema_opaque_paths`** `string[][]`

   Object paths whose contents do not take part in schema identity, one key segment per entry.
   Records that differ only inside an opaque path share one schema version. Changing this re-keys
   the collection.

- **`schema_tracked_paths`** `string[][]`

   Children of opaque paths that stay part of schema identity. Changing this re-keys the
   collection.

- **`schema_max_depth`** `number`

   How many levels of nesting take part in schema identity. Changing this re-keys the collection.

- **`schema_nulls`** [`SchemaRuleMode`](/api/collections#schema-rule-mode)

   `auto` accepts null and absent values where a mapping reads; `manual` waits for a person per
   view.

- **`schema_optional_keys`** [`SchemaRuleMode`](/api/collections#schema-rule-mode)

   `auto` lets a record with a subset of a version's keys join it; `manual` treats absence as a
   new shape.

- **`schema_detector`** [`SchemaRuleMode`](/api/collections#schema-rule-mode)

   `auto` applies a safe opaque-path proposal on its own; `manual` holds every proposal for a
   person.

- **`schema_recurrence_gap_ms`** `number`

   How long after first sight a shape must recur before its version settles, in milliseconds.

- **`schema_bulk_rows`** `number`

   A single batch carrying this many rows of a shape settles its version at once; `0` disables it.

- **`schema_stale_after_ms`** `number`

   How long a candidate version waits to recur before it goes stale, in milliseconds.

- **`schema_identity_limit`** `number`

   How many settled versions the collection may hold; past it, new shapes stay unclassified.

- **`schema_candidate_pool`** `number`

   How many candidate and stale versions the collection may hold; past it, new shapes stay
   unclassified.

- **`schema_alias_pool`** `number`

   How many exact signature hashes the collection caches for lookup.

- **`schema_writer_limit`** `number`

   How many collection files one ingest batch may open for this collection. Shapes past the
   candidate pool get the same number of files per batch, the shapes with the most rows first; the
   rows of shapes past that budget stay unclassified and cannot be linked to a version later.

- **`schema_unclassified_retention_days`** `number | null`

   How many days rows without a settled version are kept; null keeps them forever.

- **`deleted_at`** [`ISODateString | null`](/api/collections#iso-date-string)

   When the collection is scheduled to be deleted; null when it is not.

- **`created_at`** [`ISODateString`](/api/collections#iso-date-string)

- **`updated_at`** [`ISODateString`](/api/collections#iso-date-string)

### Referenced Types

#### ISODateString

`ISODateString`

An ISO 8601 date-time string returned at the JSON API boundary.

#### SchemaRuleMode

`"auto" | "manual"`

## Collection Schema Report

### Fields

- **`object`** `"collection_schema_report"`

- **`collection_id`** `string`

- **`rules`** `{ opaque_paths: string[][]; tracked_paths: string[][]; max_depth: number; }`

   The rules the report reads the collection's shapes under: the collection's own for the
   retrieve, the rules sent for a preview.

- **`shapes`** `{ before: number; after: number; }`

   How many shapes the collection holds under its current rules (`before`) and under the report's
   rules (`after`). Equal unless the report previews a rule change; the difference is what a
   re-key would merge.

- **`proposals`** [`CollectionSchemaProposal[]`](/api/collections#schema-proposal)

   Where the detector sees keys being invented, with what applying each proposal would do.

- **`keys_look_like_data`** `boolean`

   Whether every shape's top-level keys are its own, such as records keyed by a date. Those keys
   are values, and no opaque path can bring the shapes together.

- **`unclassified`** `{ shapes: number; rows: number; }`

   Shapes past the collection's pools, stored with their signature but no version, and the rows
   they carry. Nothing reads them until the pool frees or a rule brings them under a version.

### Referenced Types

#### ISODateString

`ISODateString`

An ISO 8601 date-time string returned at the JSON API boundary.

## Collection Schema Proposal

### Fields

- **`object`** `"collection_schema_proposal"`

- **`path`** `string[]`

   The path the detector proposes opaque, as key segments.

- **`tracked`** `string[]`

   Children of the path every shape carries with one type, which stay in schema identity as
   tracked paths when the proposal is applied.

- **`content_children`** `number`

   How many children of the path read as content: keys that keep being invented.

- **`read_by_mapping`** `boolean`

   Whether a bound mapping reads inside the path. Applying the proposal would leave that mapping
   unverified, so the detector never applies such a proposal on its own; a person can.

- **`shapes_after`** `number`

   How many shapes the collection would hold with this proposal applied.

### Referenced Types

#### ISODateString

`ISODateString`

An ISO 8601 date-time string returned at the JSON API boundary.

## List Collections

### Endpoint

Retrieve a list of collections for a project.

```http
GET /v1/projects/:project_id/collections
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

### Query Parameters

- **`order_by`** `string`
  Field used to order the collections. Optional. Defaults to `"name"`. Allowed values: `"created_at"`, `"name"`, `"updated_at"`.

- **`source_id`** `string`
  Filter to collections under a single source. Optional.

- **`project_id`** `string`
  Filter to collections in a single project. Optional.

- **`deleted_at`** [`NullableDateFilter`](/api/collections#nullable-date-filter)
  Filter by scheduled deletion date. Use `null` for collections that are not scheduled for deletion, `not:null` for collections that are. Optional.

- **`limit`** `number`
  Optional. Defaults to `25`. Minimum: `1`. Maximum: `200`.

- **`after`** `string`
  Cursor from `pagination.next_cursor` of a previous response. Returns the resources after that page. Optional.

- **`before`** `string`
  Cursor from `pagination.prev_cursor` of a previous response. Returns the resources before that page. Optional.

- **`sort`** `string`
  Optional. Defaults to `"asc"`. Allowed values: `"asc"`, `"desc"`.

### Response

```ts
{
  message: string;
  data: Collection[];
  status: 200;
  error: null;
  pagination: Pagination;
  endpoint: string;
}
```

### Comments

- `after` and `before` are mutually exclusive.

## Retrieve Collection

### Endpoint

Retrieve a single collection.

```http
GET /v1/projects/:project_id/collections/:collection_id
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Response

Collection retrieved

```ts
{
  message: string;
  data: Collection;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

## Create Collection

### Endpoint

Create a collection under a source.

```http
POST /v1/projects/:project_id/collections
```

**Scope:** `sources:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

### Request Body

- **`name`** `string` -- **Required**
  Display name for the collection. Minimum length: `1`. Maximum length: `128`.

- **`slug`** `string`
  Permanent identifier used in ingest URLs. Derived from `name` when omitted. Optional. Minimum length: `1`. Maximum length: `64`.

- **`source_id`** `string` -- **Required**
  Source the collection is created under.

### Response

Collection created

```ts
{
  message: string;
  data: Collection;
  status: 201;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- `source_id` is required.
- `slug` is permanent. It appears in the ingest URL, so pick it deliberately.

## Update Collection

### Endpoint

Update a collection.

```http
POST /v1/projects/:project_id/collections/:collection_id
```

**Scope:** `sources:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Request Body

- **`name`** `string`
  Display name for the collection. Optional. Minimum length: `1`. Maximum length: `128`.

- **`expected_updated_at`** [`ISODateString`](/api/collections#iso-date-string)
  Reject this update if the collection changed since this inspected update time. Optional.

- **`schema_opaque_paths`** `string[][]`
  Object paths whose contents do not take part in schema identity, each written as its key segments (`["metadata", "id"]`). A path that names an array covers its elements and their children. Records that differ only inside an opaque path share one schema version; the raw values are still stored. Changing this re-keys the collection. Optional.

- **`schema_tracked_paths`** `string[][]`
  Children of opaque paths that stay part of schema identity, written like `schema_opaque_paths`. Changing this re-keys the collection. Optional.

- **`schema_max_depth`** `integer`
  How many levels of nesting take part in schema identity. Changing this re-keys the collection. Optional. Minimum: `1`. Maximum: `10`.

- **`schema_nulls`** `string`
  `auto` accepts null and absent values at any path a mapping reads; `manual` waits for a person to approve each nullable input per view. Optional. Allowed values: `"auto"`, `"manual"`.

- **`schema_optional_keys`** `string`
  `auto` lets a record whose keys are a subset of a version's keys join that version; `manual` treats every absent key as a different shape. Optional. Allowed values: `"auto"`, `"manual"`.

- **`schema_detector`** `string`
  `auto` applies a safe opaque-path proposal from the detector on its own; `manual` holds every proposal for a person. Optional. Allowed values: `"auto"`, `"manual"`.

- **`schema_recurrence_gap_ms`** `integer`
  How long after first sight a shape must be seen again before its version settles, in milliseconds. Optional. Minimum: `60000`. Maximum: `604800000`.

- **`schema_bulk_rows`** `integer`
  A single batch carrying at least this many rows of a shape settles its version immediately. `0` disables the bulk path. Optional. Minimum: `0`. Maximum: `1000000`.

- **`schema_stale_after_ms`** `integer`
  How long a candidate version waits to be seen again before it goes stale, in milliseconds. Optional. Minimum: `3600000`. Maximum: `7776000000`.

- **`schema_identity_limit`** `integer`
  How many settled schema versions the collection may hold. Past it, new shapes stay unclassified with their signatures kept. Optional. Minimum: `1`. Maximum: `4096`.

- **`schema_candidate_pool`** `integer`
  How many candidate and stale versions the collection may hold. Past it, new shapes stay unclassified. Optional. Minimum: `1`. Maximum: `8192`.

- **`schema_alias_pool`** `integer`
  How many exact signature hashes the collection remembers for fast lookup. Past it, a new hash is still resolved but not cached. Optional. Minimum: `16`. Maximum: `100000`.

- **`schema_writer_limit`** `integer`
  How many collection files one ingest batch may open for this collection. Shapes past the candidate pool get the same number of files per batch, the shapes with the most rows first; the rows of shapes past that budget stay unclassified and cannot be linked to a version later. Optional. Minimum: `1`. Maximum: `256`.

- **`schema_unclassified_retention_days`** `integer | null`
  How many days rows without a settled schema version are kept before they are deleted. `null` keeps them forever. Optional.

- **`deleted_at`** `null`
  Send `null` to restore a collection that is scheduled for deletion. The deletion date itself is set by the delete endpoint and cannot be written here. Optional.

### Response

Collection updated

```ts
{
  message: string;
  data: Collection;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- `slug` cannot be changed after creation.
- Send `deleted_at` as `null` to restore a collection that is scheduled for deletion.

## Delete Collection

### Endpoint

Schedule a collection for deletion.

```http
DELETE /v1/projects/:project_id/collections/:collection_id
```

**Scope:** `sources:delete`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Response

```ts
{
  message: string;
  data: Collection;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- The collection is hidden immediately and permanently removed 7 days later. Restore it before then by updating it with `deleted_at` set to null.

## Force Delete Collection

### Endpoint

Delete a collection permanently, without waiting out its restore window.

```http
DELETE /v1/projects/:project_id/collections/:collection_id/force
```

**Scope:** `sources:delete`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Response

Collection queued for permanent deletion.

```ts
{
  message: string;
  data: null;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- Works on a live collection as well as one already scheduled for deletion.
- Every record, schema version and file in the collection is destroyed immediately.

## List Collection Schemas

### Endpoint

Retrieve the schema versions for a collection.

```http
GET /v1/projects/:project_id/collections/:collection_id/schemas
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Query Parameters

- **`limit`** `number`
  Maximum number of items to return. Optional. Defaults to `25`. Minimum: `1`. Maximum: `200`.

- **`after`** `string`
  Cursor from `pagination.next_cursor` of a previous response. Returns the resources after that page. Optional.

- **`before`** `string`
  Cursor from `pagination.prev_cursor` of a previous response. Returns the resources before that page. Optional.

- **`sort`** `string`
  Sort direction for the result set. Optional. Defaults to `"desc"`. Allowed values: `"asc"`, `"desc"`.

- **`state`** `string`
  Return only schema versions in this state. Optional. Allowed values: `"candidate"`, `"settled"`, `"stale"`.

### Response

```ts
{
  message: string;
  data: CollectionSchemaVersion[];
  status: 200;
  error: null;
  pagination: Pagination;
  endpoint: string;
}
```

### Comments

- `state` is `candidate` from first sight, `settled` once the shape recurs or arrives in bulk, and `stale` when a candidate is not seen again. Only a settled version is offered to views for authoring.
- `after` and `before` are mutually exclusive.

## Update Collection Schema

### Endpoint

Settle a schema version by hand, so views can author against it before its shape recurs.

```http
POST /v1/projects/:project_id/collections/:collection_id/schemas/:schema_version_id
```

**Scope:** `sources:write`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

- **`schema_version_id`** `string` -- **Required**
  Unique identifier of the schema.

### Request Body

- **`state`** `string` -- **Required**
  The only state a person sets. `settled` trusts the version for view authoring now, without waiting for its shape to recur; a settled version never goes back. Allowed values: `"settled"`.

### Response

Schema version settled

```ts
{
  message: string;
  data: CollectionSchemaVersion;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- `state` accepts only `settled`; a settled version never returns to candidate or stale.
- A version that is already settled is left as it is and responds with `request_no_update`.

## Retrieve Collection Schema Report

### Endpoint

Retrieve the detector's view of a collection: the opaque paths it proposes, whether the top-level keys are data, and how many shapes sit unclassified past the collection's pools.

```http
GET /v1/projects/:project_id/collections/:collection_id/schema_report
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Response

Collection schema report retrieved

```ts
{
  message: string;
  data: CollectionSchemaReport;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- A proposal a bound mapping reads inside is reported and never applied on its own, whatever the detector switch; apply it by adding its path to the collection's `schema_opaque_paths` and its tracked children to `schema_tracked_paths`.
- `shapes.before` and `shapes.after` are equal here; the preview is where they differ.

## Preview Collection Schema Rules

### Endpoint

Preview a rule change: the schema report the collection would read under the rules sent, with how many shapes it holds now and how many it would hold. Nothing is changed.

```http
POST /v1/projects/:project_id/collections/:collection_id/schema_preview
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Request Body

- **`schema_opaque_paths`** `string[][]`
  Opaque paths to preview, written like the collection's `schema_opaque_paths`. Omitted, the collection's current opaque paths apply. Optional.

- **`schema_tracked_paths`** `string[][]`
  Tracked paths to preview, written like the collection's `schema_tracked_paths`. Omitted, the collection's current tracked paths apply. Optional.

- **`schema_max_depth`** `integer`
  Max depth to preview. Omitted, the collection's current max depth applies. Optional. Minimum: `1`. Maximum: `10`.

### Response

Collection schema rules previewed

```ts
{
  message: string;
  data: CollectionSchemaReport;
  status: 200;
  error: null;
  pagination: null;
  endpoint: string;
}
```

### Comments

- A rule left out of the body keeps the collection's current value.
- `shapes.before` counts the shapes under the collection's current rules and `shapes.after` under the rules sent; the difference is what updating the collection with those rules would merge.

## List Collection Documents

### Endpoint

Retrieve raw documents from a collection.

```http
GET /v1/projects/:project_id/collections/:collection_id/docs
```

**Scope:** `sources:read`

### Path Parameters

- **`project_id`** `string` -- **Required**
  Unique identifier of the project.

- **`collection_id`** `string` -- **Required**
  Unique identifier of the collection.

### Response

```ts
{
  message: string;
  data: CollectionDoc[];
  status: 200;
  error: null;
  pagination: Pagination;
  endpoint: string;
}
```

### Comments

- Documents are returned exactly as they were sent, minus Tailglow's own reserved envelope.
- `after` and `before` are mutually exclusive.

