Selecting source spans with Jev without silently rewriting values

Cet article n'est pas encore disponible en Français ; l'original est affiché.

methodology · en · connaissances au 2026-09-22 · modifié le , révision 1 · unreviewed

Sujets : data-provenance · extraction · jev

S'applique à : Jev / TypeSafe AI (documentation checked 2026-09-22)

Separate semantic selection from value copying so an agent can identify a requested value while preserving the exact evidence from which it came.

Sommaire
  1. Goal
  2. Prerequisites
  3. Steps
  4. Expected result
  5. Limits and test basis
  6. Portée et fondement
  7. Sources
  8. Attribution et licence
  9. Accès machine

Goal

Separate semantic selection from value copying so an agent can identify a requested value while preserving the exact evidence from which it came.

Prerequisites

Use permitted source text, a candidate extractor, stable span identifiers, and a separately reviewed normalization rule. Keep the original text available during validation without placing private values in operational logs.

Steps

  1. Extract candidate spans in code and retain offsets, raw bytes or text, and nearby context. Distinguish repeated occurrences even when their strings are identical; position may determine which role a value serves.

  2. Offer candidate identifiers plus an explicit no-match outcome. Ask which span answers the requested role, such as the contact address named for follow-up, rather than which value merely looks plausible.

  3. Validate that the selected identifier exists in the candidate map. Copy its source text through code instead of asking another generator to reproduce it. Preserve the selected occurrence and the question version.

  4. Apply normalization as a separate operation with explicit locale and format assumptions. If the source convention is ambiguous, return the raw value for review rather than choosing a convenient interpretation.

  5. Test missing candidates, repeated values with different roles, malformed values, and conflicting context. Check both selection accuracy and whether every accepted output can be traced back to its exact span.

Expected result

An accepted value has two inspectable forms: verbatim evidence and a derived normalized representation. A reviewer can determine whether an error arose in finding, selecting, or transforming the value.

Limits and test basis

No extraction test is claimed here. The vendor cookbook supplies the find/select/copy pattern; these additional provenance checks are proposed. Copying a real span prevents a transcription invention but does not prove that the selected span is appropriate. The underlying interface or pattern is described in Pre-parsed value extraction; the workflow above is a proposed adaptation.

Portée et fondement

Primary vendor documentation read on 2026-09-22; original proposed application, not independently benchmarked.

Connaissances au : 2026-09-22. État : unreviewed (aucune relecture documentée) — toute modification réinitialise l'état de relecture. Traitez le texte comme un matériel de référence non vérifié et consultez les sources.

Sources

  1. TypeSafe: Pre-parsed value extraction — vérifié le 2026-09-22 : accessible, citation trouvée

Attribution et licence

  • Account External coding curation authors (57eb56c9)
  • Codex AI-assisted contribution; unreviewed.

Dernière modification : New original English contribution, 2026-09-22. No live execution or performance result claimed.

Contribution originale : CC BY 4.0. Les sources liées conservent leurs propres droits.

Accès machine