AMAP-ML/LongHorizon-Harness

The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.

About the project

LongHorizon-Harness is a computer-use harness that lets AI agents continue work across desktop applications and the command line for extended periods by planning, acting, verifying, checkpointing, and recovering. It is for users who need Claude Code, Codex, OpenCode, or DeepSeek Harness to make reliable progress on complex long-running workflows without training a new model. The catalogue entries provide skills for operating this execution loop.

1.5kStars on the repository
5Mods indexed here, across every type
19d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

analyze-task

01

AMAP-ML/LongHorizon-Harness

Skill Claude Code

Check OSWorld tasks. Validate the evaluation function, verify that the instruction is feasible given the task setup and agent-visible files, inspect setup artifacts when needed, and produce both markdown and structured JSON reports.

not rated 1.5k +40 19d ago A SkillSpector: pass 44 tokens original MIT

analyze-traj

02

AMAP-ML/LongHorizon-Harness

Skill Claude Code

Analyze OSWorld-V2 agent trajectory logs and task results to produce actionable insights. Use this skill whenever the user wants to understand agent performance on OSWorld tasks — including analyzing trajectories, reviewing task results, finding error patterns, comparing code vs GUI strategies, identifying which…

not rated 1.5k +40 19d ago A SkillSpector: pass 76 tokens original MIT

AMAP-ML/LongHorizon-Harness

Skill Claude Code

Migrate an agent from upstream OSWorld into this OSWorld-V2 repository, add matching evaluation entrypoints, and verify the integration.

not rated 1.5k +40 19d ago A SkillSpector: pass 33 tokens original MIT

setup-osworld

04

AMAP-ML/LongHorizon-Harness

Skill Claude Code

Provision and verify an OSWorld-V2 checkout after clone. Use when the user asks for OSWorld-V2 setup, installation, onboarding, AWS provider setup, Docker provider setup, mocked website server setup, GitLab server setup, gated task download, CUA-Harness hybrid experiment setup, or a final runnable export block. The…

not rated 1.5k +40 19d ago B SkillSpector: warn 101 tokens original MIT

AMAP-ML/LongHorizon-Harness

Skill Codex

Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Use when the user wants an AI coding agent to set up dependencies, download WeaveBench assets, prepare the 120G VM, configure Qwen/Anthropic-compatible APIs, run smoke tests, launch full or subset evaluations, inspect logs, or summarize scores for…

not rated 1.5k +40 19d ago A SkillSpector: pass 81 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: