주제: measurement
-
퇴비 온도 기록: 고정된 측정 지점과 깊이, 더미 옆의 기온, 뒤집기를 모두 이벤트로 남기기
정원 퇴비 더미나 퇴비통을 위한 관찰 프로토콜 제안입니다. 긴 탐침 온도계로 표시해 둔 측정 지점과 정해진 깊이를 정해진 일정에 따라 재고, 같은 순간 더미 옆의 기온도 함께 재며, 재료 추가·뒤집기·물 주기를 모두 이벤트 행으로 기록해, 더미의 온도가 오르고 정체되고 내려가는 과정을 그 사이에 한 조치들과 함께 읽을 수 있게 합니다. 목표 온도나 결과에 대해서는 어떠한 주장도 하지 않습니다.
-
중앙값, 백분위수, 비율에 대한 신뢰구간을 부트스트랩으로 구하기
원본 관측값에서 복원추출로 여러 번 재표본을 뽑아 매번 통계량을 계산하고, 그렇게 얻은 분포에서 구간을 읽어냅니다. 이 방법은 교과서 공식이 없는 중앙값, 백분위수, 비율, 그리고 이들의 차이에 대해서도 불확실성을 제공합니다. 방법, 재표본 횟수, 표본 크기를 함께 보고해야 하며, 소표본의 극단적인 백분위수에는 이 방법을 신뢰해서는 안 됩니다.
-
가정용 온도계를 얼음물 중탕에서 읽은 값 기록하기: 기기별 오프셋 기록
NIST가 설명하는 얼음의 녹는점 조건(증류수로 만든 잘게 부순 얼음, 위에서 아래까지 얼음물이 섞인 상태, 정해진 침지 깊이)을 따라, 가정용 온도계 각각이 명목상 0°C에서 실제로 어떤 값을 가리키는지 날짜, 준비 방법, 값이 안정되기까지 걸린 시간과 함께 기록하는, 기록 전용 프로토콜 제안입니다. 기기별 오프셋 이력을 남길 뿐, 보정 방법이나 식품 용도에 대한 지침은 제공하지 않습니다.
-
변경 사항 벤치마킹하기: 워밍업, 반복, 분산, 그리고 무엇을 보고할 것인가
시간 측정 비교는 잡음을 이겨 내야만 비로소 결과라고 부를 수 있습니다. 워크로드를 고정하고, 워밍업 실행은 버리고, 각 변형(variant)을 여러 번 번갈아 실행하고, 데이터를 보기 전에 어떤 통계량을 쓸지 정하고, 모든 수치 옆에 산포와 환경을 함께 보고해야 합니다. 실행 간 산포보다 작은 차이는 유의미한 결과가 아닙니다.
-
가정에서 씨앗 발아 비교 실험을 어떻게 설계해야 두 가정의 결과를 서로 비교할 수 있을까?
열린 질문: 연구소는 ISTA(국제종자검정협회)의 국제 종자 검정 규정에 따라 씨앗을 검정하지만, 씨앗 두 로트나 창턱 두 곳을 비교하는 가정에는 공유된 프로토콜이 없습니다. 어떤 표본 크기, 판정 기준, 기간, 조건 기록이 있어야 이런 가정 내 비교가 유의미해지고 가정끼리도 비교할 수 있게 될까요?
-
Measuring typing speed at home: a fixed-text, fixed-duration protocol with the word and error rules written down
A proposed protocol for a personal typing-speed record in which the definitions are part of the log: a standard word is defined as five characters including spaces, gross and net rates are computed by stated formulas, the keyboard, layout, software and correction setting are logged per session, three timed trials of fixed length are run on texts of a fixed kind, and the median is reported; no rate, target or improvement is claimed.
-
How much of an agent's context is tool output in real runs, and does trimming it change task success?
Open question: the MCP specification says clients should validate tool results before passing them to the model but leaves the amount to the client; in recorded agent runs, what share of tokens is tool output rather than instructions or reasoning, and does truncating, summarising or filtering tool output change task success, cost and latency?
-
A device battery health log: what phones, Windows laptops and the Linux power-supply interface report
A monthly log of the battery figures a device reports about itself: the iPhone Battery Health screen's maximum capacity relative to new, the HTML report from powercfg /batteryreport on Windows, and the charge_full and charge_full_design attributes of the Linux power-supply class; recorded raw with date, software version and events, the series shows the trend and its jumps without any charging advice.
-
Making a recipe substitution experiment comparable
A proposed protocol for documenting an ingredient substitution: fix everything except the substituted ingredient, record quantities by mass, describe the equipment and timings, and separate measured observations from preference judgements; no cooking result is asserted.
-
A grocery price log with the unit price computed from the pack: barcode identity, pack quantity, shelf price and promotion flag per observation
A proposed record-only protocol for logging grocery prices in which each observation carries the product's barcode number as its identity, the pack quantity as printed, the shelf price and any promotion or loyalty condition, and a unit price computed by the household in a stated unit, following the definition in EU Directive 98/6/EC of the unit price as the final price per kilogramme, litre, metre, square metre or cubic metre; a changed pack size becomes a new product row, and no purchasing advice or inflation figure is given.
-
An explicit 'none of these' option in every closed decision lowers an agent's wrong-action rate more than raising the confidence threshold does
For an agent that routes or classifies with a closed set of options and acts on the result, this hypothesis predicts that adding an explicit abstain option to the option set removes more wrong actions per blocked correct action than tightening a confidence threshold on the same question without such an option.
-
Household log entries written from memory at the end of the day show more rounded values than entries written at the moment of reading
Hypothesis: when a meter, scale or thermometer reading is written down hours later from memory, the recorded value is more often a round number (ending in 0 or 5, or with fewer decimals) than when it is written at the instrument; a proposed within-household test with alternating days and photographs as the reference, with no claim about which value is closer to the truth.
-
Measurement uncertainty and significant figures in technical reports
A measured value without an uncertainty is incomplete: repeat the measurement, report the mean with the standard deviation of the mean and the number of runs, round the uncertainty to one or two significant figures and the value to the same place, and say what the interval means; digits beyond the uncertainty are noise.
-
Measuring what you learned with before-and-after self-tests, and what such a comparison cannot show
A protocol for a personal pre-test and post-test around a study period: write the questions before studying, answer them blind, study, answer a parallel set after a delay and score with a fixed key; the difference is an estimate with known weaknesses, since the pre-test itself teaches, the sets may differ in difficulty, and a single learner cannot be their own control.
-
Checking a kitchen scale with coins of published mass: a repeatability, range and corner-load record
A proposed record-only protocol for a household scale: use new coins whose nominal mass the issuing mint publishes as reference masses, log readings for single coins and stacks across the scale's range, repeat placements, test the four corners and a timed hold, and keep the sheet per scale; no adjustment and no pass or fail judgement is part of it.
-
A watering and growth log for houseplants: fixed measurement points and photographs
A proposed observation log for potted plants: measurement points defined once (height from the pot rim, leaf count above a stated size, largest leaf length), watering recorded by volume or mass, position and events, and a weekly photograph under the same conditions, so that growth can be compared over time and between plants; no care advice is given.
-
Two consumer hygrometers in one room disagree less on dew point than on relative humidity
Hypothesis: two consumer temperature-humidity devices placed at different spots in one room report relative humidity values that differ mainly because their temperatures differ, so the dew points computed from each device's own pair agree more closely than the raw RH readings once the fixed inter-device offset is removed; no measurement is reported.
-
Load testing with open and closed workload models
In a closed model a fixed number of virtual users wait for each response before sending the next request, so a slowing server throttles its own load and the worst periods go unmeasured; in an open model requests arrive at a set rate regardless of completion. Choose the model from the question being asked and report it with every number.
-
Measuring home internet throughput repeatably: a fixed-path, fixed-schedule protocol
A proposed protocol for a household throughput and latency series: hold the device, the wired or wireless path, the test tool and the server constant, run three consecutive tests in three fixed daily slots for two weeks, log connection count and household activity with each run, and read RFC 6349's bandwidth-delay product and single-versus-multiple-connection points as reasons why tool settings change the number; nothing is claimed about any provider.
-
How should the reliability of an acting agent be measured when a run can succeed at its task and still cause an unwanted side effect?
Open question: benchmarks score whether the goal state was reached, and pass^k adds consistency over trials, but neither counts a run that reached the goal and also deleted a file, sent a message or spent a budget it should not have; which measures teams use for that, how they collect them, and whether they move with prompt and model changes is undocumented.
기계 판독 가능: JSON