agentic evaluation skills

27 tagged agentic evaluation, measured the same way as everything else here.

Browse within: benchmarks 27openclaw 27

InternLM/WildClawBench

Skill Claude CodeCodex

Fetches and summarizes recent arXiv and Hugging Face papers with Agentic Paper Digest. Use when the user wants a paper digest, a JSON feed of recent papers, or to run the arXiv/HF pipeline.

515 15d ago A 54 tokens original MIT

agenticmail

02

InternLM/WildClawBench

Skill Claude CodeCodex

πŸŽ€ AgenticMail β€” Full email, SMS, storage & multi-agent coordination for AI agents. 63 tools.

515 15d ago A 28 tokens original MIT

InternLM/WildClawBench

Skill Claude CodeCodex

Booking links fail for groups. SkipUp schedules meetings with 2-50 participants via email β€” one API call coordinates across timezones automatically. Also: check status, pause, resume, or cancel requests. Async only β€” does not instant-book, access calendars, or do free/busy lookups.

515 15d ago A 66 tokens original MIT