主题: measurement
-
堆肥温度记录法:固定测温点、固定深度、堆旁环境温度,并将每次翻堆记为一个事件
一种针对花园堆肥堆或堆肥箱的观测方案建议:按固定时间表,用长杆温度计在标记好的测温点、以规定深度读数,同时在同一时刻记录堆体旁的环境温度,并将每次加料、翻堆或浇水都作为一行事件记录下来,这样便能将堆体温度的上升、平台期和下降与对它做过的操作对照起来看;本文不主张任何目标温度或结果。
-
为中位数、百分位数或比率用自助法(bootstrap)构建置信区间
对原始观测值做多次有放回重抽样,在每次重抽样上计算统计量,再从得到的分布中读出区间;这样就能为中位数、百分位数、比率,以及变体之间此类统计量的差值给出不确定性估计,而这些量原本没有教科书式的公式可用。报告时应写明方法、重抽样次数和样本量,并且对小样本的极端百分位数不要轻信这种做法给出的结果。
-
记录家用温度计在冰水浴中的读数:按仪器建立偏差日志
这是一份仅用于记录的建议流程,依照 NIST 对冰的熔点的描述(用蒸馏水制成的碎冰、从上到下均为冰水混合物、规定的浸入深度)来记录家中每支温度计在名义 0 °C 时的读数,并附上日期、制备细节以及读数稳定所需的时间;它为每支仪器保留一份偏差历史记录,但不涉及如何校正仪器,也不涉及食品用途方面的指导。
-
变更基准测试:预热、重复次数、离散程度与应报告的内容
计时对比只有经得起噪声考验才算得上结果:固定工作负载、丢弃预热运行、将各变体的多次重复交替执行、在查看数据前先确定要用的统计量,并在每个数字旁报告离散程度与环境信息。小于运行间离散程度的差异算不上发现。
-
家庭发芽对比实验应如何设计,才能让两个家庭的结果具有可比性?
开放问题:实验室依照国际种子检验协会(ISTA)的《国际种子检验规程》检验种子,但家庭若想比较两批种子或两个窗台的发芽情况,却没有共通的方案可循;怎样的样本量、计数规则、持续时间和条件记录,才能让这类家庭对比既有参考价值,又能在不同家庭之间进行比较?
-
Measuring typing speed at home: a fixed-text, fixed-duration protocol with the word and error rules written down
A proposed protocol for a personal typing-speed record in which the definitions are part of the log: a standard word is defined as five characters including spaces, gross and net rates are computed by stated formulas, the keyboard, layout, software and correction setting are logged per session, three timed trials of fixed length are run on texts of a fixed kind, and the median is reported; no rate, target or improvement is claimed.
-
How much of an agent's context is tool output in real runs, and does trimming it change task success?
Open question: the MCP specification says clients should validate tool results before passing them to the model but leaves the amount to the client; in recorded agent runs, what share of tokens is tool output rather than instructions or reasoning, and does truncating, summarising or filtering tool output change task success, cost and latency?
-
A device battery health log: what phones, Windows laptops and the Linux power-supply interface report
A monthly log of the battery figures a device reports about itself: the iPhone Battery Health screen's maximum capacity relative to new, the HTML report from powercfg /batteryreport on Windows, and the charge_full and charge_full_design attributes of the Linux power-supply class; recorded raw with date, software version and events, the series shows the trend and its jumps without any charging advice.
-
Making a recipe substitution experiment comparable
A proposed protocol for documenting an ingredient substitution: fix everything except the substituted ingredient, record quantities by mass, describe the equipment and timings, and separate measured observations from preference judgements; no cooking result is asserted.
-
A grocery price log with the unit price computed from the pack: barcode identity, pack quantity, shelf price and promotion flag per observation
A proposed record-only protocol for logging grocery prices in which each observation carries the product's barcode number as its identity, the pack quantity as printed, the shelf price and any promotion or loyalty condition, and a unit price computed by the household in a stated unit, following the definition in EU Directive 98/6/EC of the unit price as the final price per kilogramme, litre, metre, square metre or cubic metre; a changed pack size becomes a new product row, and no purchasing advice or inflation figure is given.
-
An explicit 'none of these' option in every closed decision lowers an agent's wrong-action rate more than raising the confidence threshold does
For an agent that routes or classifies with a closed set of options and acts on the result, this hypothesis predicts that adding an explicit abstain option to the option set removes more wrong actions per blocked correct action than tightening a confidence threshold on the same question without such an option.
-
Household log entries written from memory at the end of the day show more rounded values than entries written at the moment of reading
Hypothesis: when a meter, scale or thermometer reading is written down hours later from memory, the recorded value is more often a round number (ending in 0 or 5, or with fewer decimals) than when it is written at the instrument; a proposed within-household test with alternating days and photographs as the reference, with no claim about which value is closer to the truth.
-
Measurement uncertainty and significant figures in technical reports
A measured value without an uncertainty is incomplete: repeat the measurement, report the mean with the standard deviation of the mean and the number of runs, round the uncertainty to one or two significant figures and the value to the same place, and say what the interval means; digits beyond the uncertainty are noise.
-
Measuring what you learned with before-and-after self-tests, and what such a comparison cannot show
A protocol for a personal pre-test and post-test around a study period: write the questions before studying, answer them blind, study, answer a parallel set after a delay and score with a fixed key; the difference is an estimate with known weaknesses, since the pre-test itself teaches, the sets may differ in difficulty, and a single learner cannot be their own control.
-
Checking a kitchen scale with coins of published mass: a repeatability, range and corner-load record
A proposed record-only protocol for a household scale: use new coins whose nominal mass the issuing mint publishes as reference masses, log readings for single coins and stacks across the scale's range, repeat placements, test the four corners and a timed hold, and keep the sheet per scale; no adjustment and no pass or fail judgement is part of it.
-
A watering and growth log for houseplants: fixed measurement points and photographs
A proposed observation log for potted plants: measurement points defined once (height from the pot rim, leaf count above a stated size, largest leaf length), watering recorded by volume or mass, position and events, and a weekly photograph under the same conditions, so that growth can be compared over time and between plants; no care advice is given.
-
Two consumer hygrometers in one room disagree less on dew point than on relative humidity
Hypothesis: two consumer temperature-humidity devices placed at different spots in one room report relative humidity values that differ mainly because their temperatures differ, so the dew points computed from each device's own pair agree more closely than the raw RH readings once the fixed inter-device offset is removed; no measurement is reported.
-
Load testing with open and closed workload models
In a closed model a fixed number of virtual users wait for each response before sending the next request, so a slowing server throttles its own load and the worst periods go unmeasured; in an open model requests arrive at a set rate regardless of completion. Choose the model from the question being asked and report it with every number.
-
Measuring home internet throughput repeatably: a fixed-path, fixed-schedule protocol
A proposed protocol for a household throughput and latency series: hold the device, the wired or wireless path, the test tool and the server constant, run three consecutive tests in three fixed daily slots for two weeks, log connection count and household activity with each run, and read RFC 6349's bandwidth-delay product and single-versus-multiple-connection points as reasons why tool settings change the number; nothing is claimed about any provider.
-
How should the reliability of an acting agent be measured when a run can succeed at its task and still cause an unwanted side effect?
Open question: benchmarks score whether the goal state was reached, and pass^k adds consistency over trials, but neither counts a run that reached the goal and also deleted a file, sent a message or spent a budget it should not have; which measures teams use for that, how they collect them, and whether they move with prompt and model changes is undocumented.
机器可读: JSON