Tema: measurement
-
Um registo de temperatura de compostagem: pontos de sonda fixos, profundidade fixa, temperatura ambiente junto à pilha e cada reviramento como evento
Um protocolo de observação proposto para uma pilha ou caixa de compostagem de jardim: um termómetro de haste longa lido em pontos de sonda marcados e a uma profundidade indicada, segundo um horário fixo, com a temperatura ambiente junto à pilha registada no mesmo momento, e cada adição, reviramento ou rega registados como uma linha de evento, de modo a que a subida, o patamar e a descida da pilha possam ser lidas em função do que lhe foi feito; não se afirma nenhuma temperatura-alvo nem resultado.
-
Construir por bootstrap um intervalo de confiança para uma mediana, percentil ou rácio
Reamostre as observações em bruto com reposição muitas vezes, calcule a estatística em cada reamostragem, e leia o intervalo a partir da distribuição resultante; isto dá uma incerteza para medianas, percentis, rácios e diferenças para os quais não existe fórmula de manual. Reporte o método, o número de reamostragens e o tamanho da amostra, e não confie no método para percentis extremos de amostras pequenas.
-
Registar a leitura de um termómetro doméstico num banho de água com gelo: um registo de desvio por instrumento
Um protocolo proposto, apenas de registo, que segue a descrição do NIST para o ponto de fusão do gelo (gelo picado feito de água destilada, uma mistura de gelo e água do topo à base, profundidade de imersão indicada) para registar o que cada termómetro doméstico marca a nominalmente 0 °C, com data, detalhes da preparação e o tempo que a leitura demorou a estabilizar; mantém um histórico de desvios (offsets) por instrumento e não dá nenhuma orientação sobre ajuste ou sobre uso alimentar.
-
Fazer benchmark de uma alteração: aquecimento, repetições, variância e o que reportar
Uma comparação de tempos só é um resultado se sobreviver ao ruído: fixe a carga de trabalho, descarte as execuções de aquecimento, intercale muitas repetições de cada variante, escolha a estatística antes de olhar para os dados, e reporte a dispersão e o ambiente junto a cada número. Uma diferença menor do que a dispersão entre execuções não é uma conclusão.
-
Como deve ser montada uma comparação doméstica de germinação de sementes para que duas casas possam comparar resultados?
Pergunta em aberto: os laboratórios testam sementes segundo as International Rules for Seed Testing da ISTA, mas as casas que comparam dois lotes de sementes, ou dois parapeitos de janela, não partilham nenhum protocolo; que tamanhos de amostra, regras de contagem, durações e registos de condições tornam essas comparações domésticas informativas e comparáveis entre casas?
-
Measuring typing speed at home: a fixed-text, fixed-duration protocol with the word and error rules written down
A proposed protocol for a personal typing-speed record in which the definitions are part of the log: a standard word is defined as five characters including spaces, gross and net rates are computed by stated formulas, the keyboard, layout, software and correction setting are logged per session, three timed trials of fixed length are run on texts of a fixed kind, and the median is reported; no rate, target or improvement is claimed.
-
How much of an agent's context is tool output in real runs, and does trimming it change task success?
Open question: the MCP specification says clients should validate tool results before passing them to the model but leaves the amount to the client; in recorded agent runs, what share of tokens is tool output rather than instructions or reasoning, and does truncating, summarising or filtering tool output change task success, cost and latency?
-
A device battery health log: what phones, Windows laptops and the Linux power-supply interface report
A monthly log of the battery figures a device reports about itself: the iPhone Battery Health screen's maximum capacity relative to new, the HTML report from powercfg /batteryreport on Windows, and the charge_full and charge_full_design attributes of the Linux power-supply class; recorded raw with date, software version and events, the series shows the trend and its jumps without any charging advice.
-
Making a recipe substitution experiment comparable
A proposed protocol for documenting an ingredient substitution: fix everything except the substituted ingredient, record quantities by mass, describe the equipment and timings, and separate measured observations from preference judgements; no cooking result is asserted.
-
A grocery price log with the unit price computed from the pack: barcode identity, pack quantity, shelf price and promotion flag per observation
A proposed record-only protocol for logging grocery prices in which each observation carries the product's barcode number as its identity, the pack quantity as printed, the shelf price and any promotion or loyalty condition, and a unit price computed by the household in a stated unit, following the definition in EU Directive 98/6/EC of the unit price as the final price per kilogramme, litre, metre, square metre or cubic metre; a changed pack size becomes a new product row, and no purchasing advice or inflation figure is given.
-
An explicit 'none of these' option in every closed decision lowers an agent's wrong-action rate more than raising the confidence threshold does
For an agent that routes or classifies with a closed set of options and acts on the result, this hypothesis predicts that adding an explicit abstain option to the option set removes more wrong actions per blocked correct action than tightening a confidence threshold on the same question without such an option.
-
Household log entries written from memory at the end of the day show more rounded values than entries written at the moment of reading
Hypothesis: when a meter, scale or thermometer reading is written down hours later from memory, the recorded value is more often a round number (ending in 0 or 5, or with fewer decimals) than when it is written at the instrument; a proposed within-household test with alternating days and photographs as the reference, with no claim about which value is closer to the truth.
-
Measurement uncertainty and significant figures in technical reports
A measured value without an uncertainty is incomplete: repeat the measurement, report the mean with the standard deviation of the mean and the number of runs, round the uncertainty to one or two significant figures and the value to the same place, and say what the interval means; digits beyond the uncertainty are noise.
-
Measuring what you learned with before-and-after self-tests, and what such a comparison cannot show
A protocol for a personal pre-test and post-test around a study period: write the questions before studying, answer them blind, study, answer a parallel set after a delay and score with a fixed key; the difference is an estimate with known weaknesses, since the pre-test itself teaches, the sets may differ in difficulty, and a single learner cannot be their own control.
-
Checking a kitchen scale with coins of published mass: a repeatability, range and corner-load record
A proposed record-only protocol for a household scale: use new coins whose nominal mass the issuing mint publishes as reference masses, log readings for single coins and stacks across the scale's range, repeat placements, test the four corners and a timed hold, and keep the sheet per scale; no adjustment and no pass or fail judgement is part of it.
-
A watering and growth log for houseplants: fixed measurement points and photographs
A proposed observation log for potted plants: measurement points defined once (height from the pot rim, leaf count above a stated size, largest leaf length), watering recorded by volume or mass, position and events, and a weekly photograph under the same conditions, so that growth can be compared over time and between plants; no care advice is given.
-
Two consumer hygrometers in one room disagree less on dew point than on relative humidity
Hypothesis: two consumer temperature-humidity devices placed at different spots in one room report relative humidity values that differ mainly because their temperatures differ, so the dew points computed from each device's own pair agree more closely than the raw RH readings once the fixed inter-device offset is removed; no measurement is reported.
-
Load testing with open and closed workload models
In a closed model a fixed number of virtual users wait for each response before sending the next request, so a slowing server throttles its own load and the worst periods go unmeasured; in an open model requests arrive at a set rate regardless of completion. Choose the model from the question being asked and report it with every number.
-
Measuring home internet throughput repeatably: a fixed-path, fixed-schedule protocol
A proposed protocol for a household throughput and latency series: hold the device, the wired or wireless path, the test tool and the server constant, run three consecutive tests in three fixed daily slots for two weeks, log connection count and household activity with each run, and read RFC 6349's bandwidth-delay product and single-versus-multiple-connection points as reasons why tool settings change the number; nothing is claimed about any provider.
-
How should the reliability of an acting agent be measured when a run can succeed at its task and still cause an unwanted side effect?
Open question: benchmarks score whether the goal state was reached, and pass^k adds consistency over trials, but neither counts a run that reached the goal and also deleted a file, sent a message or spent a budget it should not have; which measures teams use for that, how they collect them, and whether they move with prompt and model changes is undocumented.
Legível por máquina: JSON