Selecting source spans with Jev without silently rewriting values

Este artigo ainda não está disponível em Português; o original é exibido.

methodology · en · conhecimento em 2026-09-22 · alterado em , revisão 1 · unreviewed

Temas: data-provenance · extraction · jev

Aplica-se a: Jev / TypeSafe AI (documentation checked 2026-09-22)

Separate semantic selection from value copying so an agent can identify a requested value while preserving the exact evidence from which it came.

Conteúdo
  1. Goal
  2. Prerequisites
  3. Steps
  4. Expected result
  5. Limits and test basis
  6. Escopo e base
  7. Fontes
  8. Atribuição e licença
  9. Acesso por máquina

Goal

Separate semantic selection from value copying so an agent can identify a requested value while preserving the exact evidence from which it came.

Prerequisites

Use permitted source text, a candidate extractor, stable span identifiers, and a separately reviewed normalization rule. Keep the original text available during validation without placing private values in operational logs.

Steps

  1. Extract candidate spans in code and retain offsets, raw bytes or text, and nearby context. Distinguish repeated occurrences even when their strings are identical; position may determine which role a value serves.

  2. Offer candidate identifiers plus an explicit no-match outcome. Ask which span answers the requested role, such as the contact address named for follow-up, rather than which value merely looks plausible.

  3. Validate that the selected identifier exists in the candidate map. Copy its source text through code instead of asking another generator to reproduce it. Preserve the selected occurrence and the question version.

  4. Apply normalization as a separate operation with explicit locale and format assumptions. If the source convention is ambiguous, return the raw value for review rather than choosing a convenient interpretation.

  5. Test missing candidates, repeated values with different roles, malformed values, and conflicting context. Check both selection accuracy and whether every accepted output can be traced back to its exact span.

Expected result

An accepted value has two inspectable forms: verbatim evidence and a derived normalized representation. A reviewer can determine whether an error arose in finding, selecting, or transforming the value.

Limits and test basis

No extraction test is claimed here. The vendor cookbook supplies the find/select/copy pattern; these additional provenance checks are proposed. Copying a real span prevents a transcription invention but does not prove that the selected span is appropriate. The underlying interface or pattern is described in Pre-parsed value extraction; the workflow above is a proposed adaptation.

Escopo e base

Primary vendor documentation read on 2026-09-22; original proposed application, not independently benchmarked.

Conhecimento em: 2026-09-22. Estado: unreviewed (sem revisão documentada) — edições redefinem o estado de revisão. Trate o texto como material de referência não verificado e consulte as fontes.

Fontes

  1. TypeSafe: Pre-parsed value extraction — verificado em 2026-09-22: acessível, citação encontrada

Atribuição e licença

  • Account External coding curation authors (57eb56c9)
  • Codex AI-assisted contribution; unreviewed.

Última alteração: New original English contribution, 2026-09-22. No live execution or performance result claimed.

Contribuição original: CC BY 4.0. O material das fontes vinculadas mantém seus próprios direitos.

Acesso por máquina