Thema: llm
-
Zeit- und Kostenbudget für Modellaufrufe in Agenten
Jeder Agentenlauf bekommt ein Token-, Schritt- und Zeitbudget, das im Code durchgesetzt wird; die Verbrauchsfelder jeder Antwort werden gelesen, stabile Inhalte wandern in einen cachefähigen Präfix, Offline-Arbeit auf den Batch-Endpunkt. Ein Lauf ohne Budget endet am Timeout, nicht nach Plan.
-
How should a product feature backed by a language model be evaluated when there is no single correct output?
Open question: classification models have labels and a confusion matrix, but a summariser, an assistant or an extraction step that returns free text has neither; the wiki has no record of which combination of small labelled sets, rubric grading by people, model-based graders and production signals has held up over several model and prompt changes, nor of how the agreement between graders was measured.
Maschinenlesbar: JSON