Use BEFORE interpreting any surprising, clean-looking, or decision-relevant ML/AI experimental result, and before claiming a mechanism "works", "fails", "is weak", or "is exhausted". Enforces the skeptical artifact-first protocol (artifact hypotheses → invariance → mechanism-vs-metric → cautious interpretation).…
Use when designing an ML/AI experiment or metric to test whether a component, intervention, memory, probe, or feature CAUSALLY does something (e.g. "does the substrate use its input", "does memory carry information X", "does this module compute Y"). Provides the mandatory controls and ready-to-use diagnostic code…
Use when stating a conclusion from AI/ML experiments — especially a NEGATIVE result ("X doesn't work / is exhausted / hits a ceiling") or a positive capability claim. Enforces matching claim strength to evidence: distinguish not-measured vs optimization-failure vs fundamental-limit, scope to the tested regime, require…
Claude Code instructions for georgepok/local-llm-mcp-server, covering claude code project guide, what this is, repository structure, subproject relationships and dgx spark deployment.