Tag: validation

  • The Language of Surrender: Phrase Improvement Measurements

    The Language of Surrender: Phrase Improvement Measurements

    The trigger vocabulary that finds surrender statements was grown and tested: 373 mined candidates, 45 tested against ten labeled sources spanning four tiers of label quality, 9 promoted on human gold. Three measured regularities came out: precision rises with phrase length (0.603 at two words to 0.771 at five, over 134,600 decided spans); first-person forms separate self-acknowledged limitations from reviewer criticism (0.955 versus 0.000 for the same phrase family); and mining scores do not predict trigger quality. Every number is a measurement with its label tier stated.

  • The Language of Surrender: Larger Model Measurements

    The Language of Surrender: Larger Model Measurements

    Three models one size class up — Qwen3-14B, Phi-4, and Mistral-Small-24B at Q4_K_M quantization — measured on the same 1,505-sentence limitation gold set, same prompt, same harness as the earlier six-model comparison. Qwen3-14B scores F1 0.8272 as a single model, above the previous two-model union (0.8143) and the published rule-system reference (0.800). VRAM at load: 10.6 GB of 16.3 GB. Every number is a full-split measurement.

  • The Language of Surrender: Model Selection Observations

    The Language of Surrender: Model Selection Observations

    Scientific papers state their own defeats in words: a sample that could not be enrolled, a computation that could not be afforded, data that could not be obtained. A phrase list finds those statements with precision 0.23. Judging each match in context — with rules, two trained classifiers, and a pair of 8-billion-parameter models on two 16GB GPUs — raises the measured operating point to a 0.945-precision auto-admit gate and a two-judge F1 of 0.814, beating the published rule-based reference on the same gold set. Every number is a full-split measurement.

  • Extracting Value from Existing Structure: Leveraging Past Decisions to Predict Future Ones

    Extracting Value from Existing Structure: Leveraging Past Decisions to Predict Future Ones

    Completed work is labeled data. Every archive of past decisions pairs an input with a judgment a person already made, and an automated procedure can be replayed against that record to measure how often it agrees. The result is an accuracy figure, a failure-mode inventory, and a calibrated boundary for where automation can be trusted — obtained before the procedure touches live work.