SignalQL v0.5 — context bundles
Status: Draft with reference implementation. Part of the v0.5 specification. Builds on the object model and provenance.
v0.4 introduced CONTEXT FOR entity(…): a bounded subgraph around an entity. v0.5 adds CONTEXT FOR "task": the smallest relevant, current, trusted, provenance-preserving set of information for a task, within a token budget — without the caller needing to know whether that information began as a file, a graph node, a message, a database row, or a generated artifact.
CONTEXT FOR "design SignalQL object storage"
USING PROJECT "SignalQL"
BUDGET 32000 TOKENS
PREFER CURRENT, VERIFIED
WITH EVIDENCEThe v0.4 form is unchanged. The two are distinguished by what follows FOR: a string literal selects the v0.5 form.
Grammar (EBNF, additive)
context_task ::= "CONTEXT" "FOR" string_literal context_clause*
context_clause ::= "USING" "PROJECT" string_literal
| "BUDGET" integer "TOKENS"
| "PREFER" prefer_item ( "," prefer_item )*
| "EXCLUDE" exclude_item ( "," exclude_item )*
| "MAX" ( "OBJECTS" integer | "DEPTH" integer | "AGE" integer time_unit )
| "WITH" "EVIDENCE"
| as_of_clause
prefer_item ::= "CURRENT" | trust_level | "REPRESENTATION" rep_name
exclude_item ::= "UNTRUSTED" | "STALE" | "ROLE" role ( "FROM" role )?
trust_level ::= "UNTRUSTED" | "UNKNOWN" | "INFERRED" | "VERIFIED" | "AUTHORITATIVE"
role ::= "DATA" | "INSTRUCTION" | "PROMPT" | "POLICY" | "EXECUTABLE"
| "SECRET" | "UNTRUSTED_EXTERNAL"
explain_context ::= "EXPLAIN" "CONTEXT" context_uri
find_object ::= "FIND" "OBJECT" ( "(" name ")" )?
( "ABOUT" string_literal ( "USING" "REPRESENTATION" rep_name )? )?
( "WHERE" object_predicate )? as_of_clause? limit_clause?Clauses MAY appear in any order. Each appears at most once, except PREFER and EXCLUDE, whose items merge. The canonical form orders clauses as listed above and items in vocabulary order, so equivalent queries have one canonical text.
EXPLAIN is defined only for a bundle address. EXPLAIN <query> remains reserved for execution-plan explain.
Trust, origin, validation, and roles
Provenance says where information came from. Trust says how far to rely on it. They are separate, and v0.5 exposes both without defining what they mean.
| Field | Values (in order) |
|---|---|
origin | human, model, external, system, import, unknown |
trust | untrusted < unknown < inferred < verified < authoritative |
validation | none, automated, human |
role | data, instruction, prompt, policy, executable, secret, untrusted_external |
SignalQL does not define what "authoritative" means universally. A domain decides what earns each value; the language exposes the metadata and the order of trust levels, which is used only by PREFER and EXCLUDE.
Roles are recorded per revision and MAY be refined per representation; a representation's effective roles are the union of both. Roles are advisory metadata: EXCLUDE ROLE INSTRUCTION is only as good as the labelling. Every bundle item carries its effective roles, origin, and trust so that a consumer can fence untrusted content rather than infer it.
Semantics
Let T be the AS OF time, or the evaluation time when absent.
Scope
The candidates are live objects (not tombstoned at T) at their latest revision as of T.
USING PROJECT "p"keeps objects whoseprojectequalsp.projectis an opaque partition label.MAX AGE n unitkeeps objects whose selected revision was created within that span beforeT. A month is 30 days. WithoutAS OFthis depends on the clock, so the bundle reportsevaluated_at.
EXCLUDE is a hard filter
Applied before relevance. An excluded object is never included and is not a bridge for expansion.
| Item | Excludes |
|---|---|
UNTRUSTED | objects with trust = untrusted |
STALE | objects whose dependency state is stale or invalid |
ROLE r | representations whose effective roles include r |
ROLE r FROM s | representations whose effective roles include both r and s |
Content with the secret role MUST NOT be packed and MUST NOT be read for relevance. There is no opt-in in v0.5.
An object with no representation left to pack is excluded.
PREFER is a soft preference
PREFER changes rank. It never changes which objects are eligible.
| Item | Raises objects that are… |
|---|---|
CURRENT | in dependency state current |
| a trust level | at or above that level |
REPRESENTATION name | — (it steers which representation is packed) |
Selection
Context construction MAY optimize over semantic relevance, graph distance, recency, authority, confidence, revision status, trust, token cost, duplication, source diversity, and permissions. The algorithm is implementation-defined. What the language fixes is the meaning of the inputs above and of the output below, and these requirements:
EXCLUDE, scope, and the secret rule are never violated.- The sum of item token estimates MUST NOT exceed
BUDGET. Implementations MAY return less. - At most one representation of an object is included.
- Token counts MAY be estimates; tokenizer and model family affect them. Each item names the estimator that produced its count.
- The bundle names the relevance scorer used.
- For a given scorer and set of estimators, the same query over the same revisions MUST produce the same bundle.
Reference algorithm (lexical-v1)
The reference implementation is deterministic and inspectable. It is not required of other implementations.
- Terms. Lowercase the task, split on anything that is not a letter or digit, drop tokens shorter than two characters and the stopwords
a an and are as at be by for from in into is it of on or that the this to with, and de-duplicate. - Relevance. For each term, take the heaviest field that contains it as a whole token: title 3, the
summaryrepresentation 2, any other text representation 1. The score isfloor(sum × 1 000 000 / (3 × terms)). It is presence-based with no corpus statistics, so it is stable as a store grows. - Seeds are eligible objects scoring above zero.
- Expansion. A seed's score halves per hop over derivations and graph edges, in either direction, up to
MAX DEPTH(default 1), passing through entities and eligible objects. An object's relevance is the larger of its own score and the best score that reaches it. - Preference. Each satisfied
PREFERterm multiplies the score by 5/4, rounded down. - Rank by score descending, then dependency state (
currentfirst), then fewer hops, then object id by code unit. Keep the firstMAX OBJECTS(default 20). - Pack, pass 1 (coverage). In rank order, give each object the preferred representation if one was named, exists, and fits; otherwise its smallest representation that fits. An object for which nothing fits is left out.
- Pack, pass 2 (depth). Unless
PREFER REPRESENTATIONwas given, in rank order upgrade each object to its largest representation that still fits.
Only text representations (text/*, application/json) are packed. A representation without a stored token estimate is estimated as ceil(characters / 4) (chars-div-4-v1).
A different relevance source — an embedding index, say — plugs in as a relevance provider that scores candidates for a task. The provider's name is recorded as the bundle's scorer and is part of the bundle address.
Context Bundle
context://ctx_d73247c8d934cd364f4d22f2c10b3cea| Field | Meaning |
|---|---|
id | the bundle address |
task | the task text |
items[] | the selected representations, in rank order |
evidence[] | present with WITH EVIDENCE |
metadata | canonical_query, budget {limit, used}, scorer, estimators, as_of, evaluated_at, considered, excluded (counts by reason), truncated |
Each item carries its own provenance:
| Field | Meaning |
|---|---|
object | source object and revision, object://id@rev |
representation, media_type | which form was included |
tokens, estimator | token contribution and how it was counted |
score | rank score |
why[] | why it was included: lexical_match {terms}, graph_expansion {from, via, hops}, preferred {term} |
path[] | retrieval path from the seed to this object |
state, origin, trust, validation, roles | status and safety metadata |
content / content_ref, content_hash | the content, or where to fetch it |
WITH EVIDENCE adds references, which do not count against the budget:
- for each item, the derivations recorded from its revision — relation, source
object://id@rev, the source's latest revision, and the link state; - for each item, the entities it is linked to.
Addressing
A bundle id is content-addressed: context://ctx_ followed by the first 32 hexadecimal characters of the SHA-256 of the canonical JSON of
{ v: "0.5", q: <canonical query>, as_of, scorer, estimators,
items: [ { o: object id, r: revision, rep: representation, t: tokens, h: content hash } … ] }So the same query over the same revisions has the same address, and a new revision of any included object changes it. Addresses are comparable only for the same scorer and estimators.
This keeps the language read-only: CONTEXT FOR computes a value; it changes nothing. An address cannot be turned back into a bundle, so resolving one requires a context store — a runtime cache keyed by address, in which the first record written for an address wins. The store is the context_persistence capability; without it a bundle can be computed but not looked up later.
Recording which bundle a model run used gives full reproducibility:
sources → context://abc123 → model execution → object://xyzEXPLAIN CONTEXT
EXPLAIN CONTEXT context://ctx_d73247c8d934cd364f4d22f2c10b3ceaReturns, for a stored bundle:
| Field | Meaning |
|---|---|
context, canonical_query, as_of, computed_at, scorer, estimators, budget | how it was computed |
phases | counts: candidates, seeds, expanded, rejected, packed |
items[] | each item without content, plus matched terms and score_breakdown {lexical, propagated, boosts} |
rejected[] | every object that was considered and left out, with exactly one reason |
Rejection reasons: out_of_scope, max_age, excluded_secret, excluded_trust, excluded_stale, excluded_role, no_packable_representation, max_objects, budget. Rejected entries carry an address and a reason only.
An address that is not in the store is an unknown_context capability error.
FIND OBJECT ABOUT
ABOUT ranks objects by relevance to a text, most relevant first, returning only those scoring above zero. USING REPRESENTATION restricts relevance to one representation's content.
FIND OBJECT ABOUT "pricing strategy" USING REPRESENTATION "transcript"Columns are those of FIND OBJECT plus score.
Limits
BUDGET 1–10 000 000 tokens · MAX OBJECTS 1–500 (default 20) · MAX DEPTH 0–16 (default 1).