Query
Filter records server-side with structured query documents. Supports nested body fields, logical combinators, pagination, and an EXPLAIN plan so you know when a filter hits an index.
Overview
Query lets you ask the workspace: "give me every record where X", with X expressed as a structured filter document. The filter runs server-side — nothing streams over the wire except the matches. This is the right tool for any query that touches more than one thread, any filter over body shape (body.kind, body.priority, body._refs.*), or any workload big enough that client-side filtering wastes bandwidth.
Two entry points:
- HTTP:
POST /v1/records/querywith a JSON body describing the filter. - SDK:
client.richQuery({ filter, ... })(TypeScript) orawait client.query(filter, ...)(Python).
This guide covers both surfaces, the filter AST, and how to use the explain flag to verify an index is being used.
Quick Start
# All INTEND records on task threads from the dev agent, newest first.
curl -s -X POST http://localhost:9100/v1/records/query \
-H 'content-type: application/json' \
-d '{
"filter": {
"act": "INTEND",
"actor": "did:sync:agent:dev",
"body.kind": { "$in": ["core.task.record", "core.task.record.v1"] }
},
"sort": { "clock": -1 },
"limit": 20
}'The same query from TypeScript:
import { Client, Identity } from "@syncropel/sdk";
const client = new Client({
endpoint: "http://localhost:9100",
identity: Identity.static("did:sync:user:alice"),
});
const result = await client.richQuery({
filter: {
act: "INTEND",
actor: "did:sync:agent:dev",
"body.kind": { $in: ["core.task.record", "core.task.record.v1"] },
},
sort: { clock: -1 },
limit: 20,
});
for (const rec of result.records) {
console.log(rec.id, rec.body);
}And Python:
from syncropel import Client, Identity
client = Client(
endpoint="http://localhost:9100",
identity=Identity.static("did:sync:user:alice"),
)
records = await client.query(
filter={
"act": "INTEND",
"actor": "did:sync:agent:dev",
"body.kind": {"$in": ["core.task.record", "core.task.record.v1"]},
},
sort={"clock": -1},
limit=20,
)
for rec in records:
print(rec["id"], rec["body"])Filter Grammar
A filter is a JSON object. Top-level keys are AND-combined. Values are either scalars (sugar for $eq) or operator documents.
Comparison operators
| Operator | Semantics | Example |
|---|---|---|
$eq | field equals value | { "act": { "$eq": "KNOW" } } |
$ne | field differs from value | { "body.status": { "$ne": "cancelled" } } |
$gt | greater than | { "clock": { "$gt": 100 } } |
$gte | greater than or equal | { "clock": { "$gte": 100 } } |
$lt | less than | { "clock": { "$lt": 1000 } } |
$lte | less than or equal | { "clock": { "$lte": 1000 } } |
$in | field in list | { "actor": { "$in": ["did:...", "did:..."] } } |
$nin | field not in list | { "act": { "$nin": ["DO", "CALL"] } } |
$like | SQL LIKE pattern | { "body.title": { "$like": "music/%" } } |
$ilike | case-insensitive LIKE | { "body.title": { "$ilike": "%LOVE%" } } |
$exists | field present (true) or absent (false) | { "body.awaits": { "$exists": true } } |
A scalar value — { "act": "KNOW" } — is equivalent to { "act": { "$eq": "KNOW" } }.
Logical combinators
| Operator | Semantics | Example |
|---|---|---|
$and | all children must match | { "$and": [{ "act": "KNOW" }, { "clock": { "$gt": 0 } }] } |
$or | any child matches | { "$or": [{ "act": "DO" }, { "act": "CALL" }] } |
$not | child must not match | { "$not": { "actor": "did:sync:system:engine" } } |
Top-level fields are AND-combined implicitly, so $and is usually only needed inside $or/$not.
Allowed field paths
The parser restricts paths to a whitelist — callers cannot reach into private columns (sig, canonical_json) or invent paths that bypass the JSON1 layer.
- Top-level columns:
id,thread,actor,act,clock,data_type,namespace,created_at - Body paths:
body.<segment>[.<segment>...]where each segment matches[A-Za-z0-9_-]+
So body.kind, body.priority, and body._refs.music_artist all work; body alone (without a sub-path) is rejected.
Request Shape
{
"filter": { /* AST described above */ },
"thread": "th_optional_scope",
"sort": { "clock": -1 },
"limit": 100,
"offset": 0,
"explain": false
}thread— optional thread-scope. Shorthand for adding{ "thread": "th_..." }to the filter. Uses the thread index directly.sort— single-key document; value is1for ascending,-1for descending. Example:{ "clock": -1 }(newest first),{ "created_at": 1 }(oldest first),{ "body.priority": -1 }(highest priority first — but see Indexed Field Registry for why you'll want to declare body fields first).limit: default 100, clamped to 1000. Asking for more returns 1000 rows andcapped: true; readcappedto know a page is partial.matched_totalis disclosure-aware (it is served only when nothing was withheld from you, and otherwise counts the rows you may see), so it is a count over your own view rather than over the workspace, and it is still not a substitute forcappedwhen deciding whether you have the whole answer.offset— for pagination. Prefer{ "clock": { "$gt": last_clock } }as a cursor once you have one — cursor-style pagination is cheaper than deep offsets.explain— whentrue, the response includes aplandescribing which fields used a top-level column (indexed), and which fields fell to the body-scan path.
The same grammar, from inside an actor
A work-loop actor reaches this filter language through its query_thread
tool rather than over HTTP, and it is the same grammar running through the
same governed door:
{"thread": "th_library", "filter": {"body.kind": "track.v1", "body.labeled": false}, "count_only": true}count_only answers how many without returning the rows, which is how an
actor counts a large thread inside its budget instead of paging it into
context. Because it is the same door, a sandboxed run takes the identical
path and discloses exactly what its credential could read anyway.
Two refusals are worth knowing. A field path that is not addressable comes
back named (unknown field path: kind — body fields are body.<path>), and
it is named the same way on both transports rather than degrading into a
storage error in the sandbox. And a filter that matched nothing is reported
differently from a thread the run cannot see: the first says the filter is
valid and matched nothing here, the second says the thread may hold records
outside this run's view. Neither is proof of absence, and an actor that
confuses them will state one as the other.
Using explain
The explain flag is how you verify a query is using an index before running it against a large corpus:
curl -s -X POST http://localhost:9100/v1/records/query \
-H 'content-type: application/json' \
-d '{
"filter": { "thread": "th_abc...", "body.kind": "music.catalog.track" },
"limit": 1,
"explain": true
}' | jq '.plan'Example response:
{
"indexed_fields": ["thread"],
"unindexed_fields": ["body.kind"]
}The plan says body.kind was matched by scanning record bodies, a full scan without the right index. To upgrade that to an expression index scan, declare a body-kind manifest:
spl config add-body-kind-manifest \
--kind music.catalog.track \
--indexed-field body.kind \
--indexed-field body.titleAfter the workspace reloads config, queries on body.kind are served from the new index. The plan's unindexed_fields list still names body.kind, because that list is a conservative "this was a body.* path" marker rather than a statement about which index was chosen; measure the query time to confirm.
Patterns
"Most recent by thread"
{
"thread": "th_abc...",
"sort": { "clock": -1 },
"limit": 50
}The thread shortcut hits the primary index directly. Cheap regardless of corpus size.
"All tasks assigned to a specific agent"
{
"filter": {
"body.kind": { "$in": ["core.task.record", "core.task.record.v1"] },
"body.assigned_to": "did:sync:agent:dev"
},
"sort": { "clock": -1 }
}Declare body.kind + body.assigned_to in a manifest to avoid the body scan once your task log exceeds a few thousand records.
"Records produced in the last hour"
{
"filter": {
"created_at": { "$gte": 1719830400 }
},
"sort": { "created_at": -1 },
"limit": 100
}created_at is a top-level column; no manifest needed.
"Reference lookup"
{
"filter": {
"body._refs.music_artist": "spotify:artist:4Z8W4fKeB5YxbusRsdQVPb"
}
}References stored via Ref.* constructors in the SDK live at body._refs.<entity>. Declare each reference path you query in a manifest for speed.
"Pending decision-gate proposals"
{
"filter": {
"act": "KNOW",
"body.awaits": "actor_decision"
},
"sort": { "clock": -1 }
}Limits & Failure Modes
- Filter depth: parser handles arbitrarily nested
$and/$or, but the generated SQL complexity grows with nesting. Keep filters under ~10 levels for predictability. - Type mismatches: comparing a body field that is sometimes a string and sometimes an object returns no match; the comparison uses the stored JSON type, so
$.x = "foo"will not match when the stored$.xis an object. Use$exists: trueif you only need to check presence. - No joins: the filter operates on one record at a time. For "records where another record referencing them exists", compose two queries client-side.
- Fail-open SDK transport: both SDKs return an empty record list (not an exception) on network/5xx failures. Check
result.records.length— zero is ambiguous between "no matches" and "transport failed". Use theonEmithook or the raw HTTP call if you need to distinguish.
A query can be a record (core.query.v1)
Everything above is the query mechanism — a filter you send and a result you read. When a query is worth keeping — a view you return to, a scope you want named, a standing question — it graduates to a record: the body kind core.query.v1.
The same primitive spans the whole lifecycle. A one-off filter, a saved View, and a standing (live) query are not different kinds — they are one core.query.v1 record at different stages, folded from the th_queries thread by the query_shelf fold. The record carries:
pipeline— the bounded filter/facet/sort/window (the scope algebra, compiled to a kernel filter).name— required once the query issaved.saved/live— lifecycle flags:savedpromotes a typed-once filter to a durable View;livemarks it a standing query.provenance— where it came from:source_thread(the conversation that shaped it),trigger_record,derived_from(a prior query it was forked from).
The design maxim is "a query is a place, not a route" — a saved View has a durable identity and provenance you can revisit and share, not an ephemeral request you re-type. Queries are surfaced and promoted in the workspace (the Studio scope bar); there is no CLI verb for them today — you emit a core.query.v1 record like any other, and the query_shelf fold derives which rung of the ladder each query stands on.
core.query.v1 is distinct from core.graph.query.v1 (a federation traversal envelope — an INTEND to walk a peer's graph). A similar word for a different thing; the two are not interchangeable.
Next Steps
- Body-Kind Manifests — declare which body fields to index so rich-query filter predicates scan instead of sort.
- CEL Expressions — when the filter needs to compose with engine-side logic (triggers, routing, preconditions), CEL is the right layer.
- TypeScript SDK, Python SDK — full client reference.
Semantic Search
Free-text search over the record log. Embeds the query through a configured provider, ranks records by cosine similarity, and returns the top K. Envelope filters (thread, actor, kind) narrow the result after ranking so near-misses don't crowd out the best answer.
Projections
Schema, validators, and a markdown-subset parser for the Syncropel Rendering Protocol (SRP) — the declarative document format for query-driven and AI-generated UI.