AlligatorC0der/conkurrence

AI evaluation toolkit — measure inter-rater agreement across multiple LLM providers

0Stars on the repository
3Mods indexed here, across every type
4mo agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

conkurrence

01

AlligatorC0der/conkurrence

Plugin Claude Code

AI evaluation toolkit — measure inter-rater agreement across multiple LLM providers using Fleiss' kappa and Kendall's W.

0 4mo ago A tokens not measured

conkurrence

02

AlligatorC0der/conkurrence

Plugin Claude Code

Statistically validated consensus measurement for AI evaluation pipelines. Uses multiple AI models as independent raters, measures inter-rater reliability with Fleiss' kappa and bootstrap confidence intervals, and routes contested items to human experts.

0 4mo ago A tokens not measured