Skip to content

SignalQL v0.5 — context bundles ​

Status: Draft with reference implementation. Part of the v0.5 specification. Builds on the object model and provenance.

v0.4 introduced CONTEXT FOR entity(…): a bounded subgraph around an entity. v0.5 adds CONTEXT FOR "task": the smallest relevant, current, trusted, provenance-preserving set of information for a task, within a token budget — without the caller needing to know whether that information began as a file, a graph node, a message, a database row, or a generated artifact.

signalql
CONTEXT FOR "design SignalQL object storage"
USING PROJECT "SignalQL"
BUDGET 32000 TOKENS
PREFER CURRENT, VERIFIED
WITH EVIDENCE

The v0.4 form is unchanged. The two are distinguished by what follows FOR: a string literal selects the v0.5 form.

Grammar (EBNF, additive) ​

ebnf
context_task   ::= "CONTEXT" "FOR" string_literal context_clause*
context_clause ::= "USING" "PROJECT" string_literal
                 | "BUDGET" integer "TOKENS"
                 | "PREFER" prefer_item ( "," prefer_item )*
                 | "EXCLUDE" exclude_item ( "," exclude_item )*
                 | "MAX" ( "OBJECTS" integer | "DEPTH" integer | "AGE" integer time_unit )
                 | "WITH" "EVIDENCE"
                 | as_of_clause
prefer_item    ::= "CURRENT" | trust_level | "REPRESENTATION" rep_name
exclude_item   ::= "UNTRUSTED" | "STALE" | "ROLE" role ( "FROM" role )?
trust_level    ::= "UNTRUSTED" | "UNKNOWN" | "INFERRED" | "VERIFIED" | "AUTHORITATIVE"
role           ::= "DATA" | "INSTRUCTION" | "PROMPT" | "POLICY" | "EXECUTABLE"
                 | "SECRET" | "UNTRUSTED_EXTERNAL"

explain_context ::= "EXPLAIN" "CONTEXT" context_uri

find_object    ::= "FIND" "OBJECT" ( "(" name ")" )?
                   ( "ABOUT" string_literal ( "USING" "REPRESENTATION" rep_name )? )?
                   ( "WHERE" object_predicate )? as_of_clause? limit_clause?

Clauses MAY appear in any order. Each appears at most once, except PREFER and EXCLUDE, whose items merge. The canonical form orders clauses as listed above and items in vocabulary order, so equivalent queries have one canonical text.

EXPLAIN is defined only for a bundle address. EXPLAIN <query> remains reserved for execution-plan explain.

Trust, origin, validation, and roles ​

Provenance says where information came from. Trust says how far to rely on it. They are separate, and v0.5 exposes both without defining what they mean.

FieldValues (in order)
originhuman, model, external, system, import, unknown
trustuntrusted < unknown < inferred < verified < authoritative
validationnone, automated, human
roledata, instruction, prompt, policy, executable, secret, untrusted_external

SignalQL does not define what "authoritative" means universally. A domain decides what earns each value; the language exposes the metadata and the order of trust levels, which is used only by PREFER and EXCLUDE.

Roles are recorded per revision and MAY be refined per representation; a representation's effective roles are the union of both. Roles are advisory metadata: EXCLUDE ROLE INSTRUCTION is only as good as the labelling. Every bundle item carries its effective roles, origin, and trust so that a consumer can fence untrusted content rather than infer it.

Semantics ​

Let T be the AS OF time, or the evaluation time when absent.

Scope ​

The candidates are live objects (not tombstoned at T) at their latest revision as of T.

  • USING PROJECT "p" keeps objects whose project equals p. project is an opaque partition label.
  • MAX AGE n unit keeps objects whose selected revision was created within that span before T. A month is 30 days. Without AS OF this depends on the clock, so the bundle reports evaluated_at.

EXCLUDE is a hard filter ​

Applied before relevance. An excluded object is never included and is not a bridge for expansion.

ItemExcludes
UNTRUSTEDobjects with trust = untrusted
STALEobjects whose dependency state is stale or invalid
ROLE rrepresentations whose effective roles include r
ROLE r FROM srepresentations whose effective roles include both r and s

Content with the secret role MUST NOT be packed and MUST NOT be read for relevance. There is no opt-in in v0.5.

An object with no representation left to pack is excluded.

PREFER is a soft preference ​

PREFER changes rank. It never changes which objects are eligible.

ItemRaises objects that are…
CURRENTin dependency state current
a trust levelat or above that level
REPRESENTATION name— (it steers which representation is packed)

Selection ​

Context construction MAY optimize over semantic relevance, graph distance, recency, authority, confidence, revision status, trust, token cost, duplication, source diversity, and permissions. The algorithm is implementation-defined. What the language fixes is the meaning of the inputs above and of the output below, and these requirements:

  1. EXCLUDE, scope, and the secret rule are never violated.
  2. The sum of item token estimates MUST NOT exceed BUDGET. Implementations MAY return less.
  3. At most one representation of an object is included.
  4. Token counts MAY be estimates; tokenizer and model family affect them. Each item names the estimator that produced its count.
  5. The bundle names the relevance scorer used.
  6. For a given scorer and set of estimators, the same query over the same revisions MUST produce the same bundle.

Reference algorithm (lexical-v1) ​

The reference implementation is deterministic and inspectable. It is not required of other implementations.

  1. Terms. Lowercase the task, split on anything that is not a letter or digit, drop tokens shorter than two characters and the stopwords a an and are as at be by for from in into is it of on or that the this to with, and de-duplicate.
  2. Relevance. For each term, take the heaviest field that contains it as a whole token: title 3, the summary representation 2, any other text representation 1. The score is floor(sum × 1 000 000 / (3 × terms)). It is presence-based with no corpus statistics, so it is stable as a store grows.
  3. Seeds are eligible objects scoring above zero.
  4. Expansion. A seed's score halves per hop over derivations and graph edges, in either direction, up to MAX DEPTH (default 1), passing through entities and eligible objects. An object's relevance is the larger of its own score and the best score that reaches it.
  5. Preference. Each satisfied PREFER term multiplies the score by 5/4, rounded down.
  6. Rank by score descending, then dependency state (current first), then fewer hops, then object id by code unit. Keep the first MAX OBJECTS (default 20).
  7. Pack, pass 1 (coverage). In rank order, give each object the preferred representation if one was named, exists, and fits; otherwise its smallest representation that fits. An object for which nothing fits is left out.
  8. Pack, pass 2 (depth). Unless PREFER REPRESENTATION was given, in rank order upgrade each object to its largest representation that still fits.

Only text representations (text/*, application/json) are packed. A representation without a stored token estimate is estimated as ceil(characters / 4) (chars-div-4-v1).

A different relevance source — an embedding index, say — plugs in as a relevance provider that scores candidates for a task. The provider's name is recorded as the bundle's scorer and is part of the bundle address.

Context Bundle ​

context://ctx_d73247c8d934cd364f4d22f2c10b3cea
FieldMeaning
idthe bundle address
taskthe task text
items[]the selected representations, in rank order
evidence[]present with WITH EVIDENCE
metadatacanonical_query, budget {limit, used}, scorer, estimators, as_of, evaluated_at, considered, excluded (counts by reason), truncated

Each item carries its own provenance:

FieldMeaning
objectsource object and revision, object://id@rev
representation, media_typewhich form was included
tokens, estimatortoken contribution and how it was counted
scorerank score
why[]why it was included: lexical_match {terms}, graph_expansion {from, via, hops}, preferred {term}
path[]retrieval path from the seed to this object
state, origin, trust, validation, rolesstatus and safety metadata
content / content_ref, content_hashthe content, or where to fetch it

WITH EVIDENCE adds references, which do not count against the budget:

  • for each item, the derivations recorded from its revision — relation, source object://id@rev, the source's latest revision, and the link state;
  • for each item, the entities it is linked to.

Addressing ​

A bundle id is content-addressed: context://ctx_ followed by the first 32 hexadecimal characters of the SHA-256 of the canonical JSON of

{ v: "0.5", q: <canonical query>, as_of, scorer, estimators,
  items: [ { o: object id, r: revision, rep: representation, t: tokens, h: content hash } … ] }

So the same query over the same revisions has the same address, and a new revision of any included object changes it. Addresses are comparable only for the same scorer and estimators.

This keeps the language read-only: CONTEXT FOR computes a value; it changes nothing. An address cannot be turned back into a bundle, so resolving one requires a context store — a runtime cache keyed by address, in which the first record written for an address wins. The store is the context_persistence capability; without it a bundle can be computed but not looked up later.

Recording which bundle a model run used gives full reproducibility:

sources → context://abc123 → model execution → object://xyz

EXPLAIN CONTEXT ​

signalql
EXPLAIN CONTEXT context://ctx_d73247c8d934cd364f4d22f2c10b3cea

Returns, for a stored bundle:

FieldMeaning
context, canonical_query, as_of, computed_at, scorer, estimators, budgethow it was computed
phasescounts: candidates, seeds, expanded, rejected, packed
items[]each item without content, plus matched terms and score_breakdown {lexical, propagated, boosts}
rejected[]every object that was considered and left out, with exactly one reason

Rejection reasons: out_of_scope, max_age, excluded_secret, excluded_trust, excluded_stale, excluded_role, no_packable_representation, max_objects, budget. Rejected entries carry an address and a reason only.

An address that is not in the store is an unknown_context capability error.

FIND OBJECT ABOUT ​

ABOUT ranks objects by relevance to a text, most relevant first, returning only those scoring above zero. USING REPRESENTATION restricts relevance to one representation's content.

signalql
FIND OBJECT ABOUT "pricing strategy" USING REPRESENTATION "transcript"

Columns are those of FIND OBJECT plus score.

Limits ​

BUDGET 1–10 000 000 tokens · MAX OBJECTS 1–500 (default 20) · MAX DEPTH 0–16 (default 1).