Skip to main content
Back to Categories
Category Suite: Autonomous

Agentic Tasks Evaluation Benchmark

Multi-step refactoring, execution planning, tool use.

Category Prompts & Benchmark Tasks (10)Zero-Shot & Multi-Turn Evals
#1 Repository Schema Migration Script
Difficulty: Hard

Write a multi-step execution plan and TypeScript migration script to split monolithic user tables into tenant-scoped normalized models.

#2 Autonomous Dependency Vulnerability Audit
Difficulty: Hard

Create a multi-step CLI agent script that parses package.json files, queries vulnerability databases, evaluates breaking API changes, and auto-generates pull request patches.

#3 Multi-Agent Task Orchestrator & Task Queue
Difficulty: Hard

Implement an asynchronous multi-agent task runner with dependency DAG resolution, retries, exponential backoff, worker pool concurrency limits, and execution logging.

#4 Automated API Documentation Generator
Difficulty: Medium

Write a static analysis agent script that parses AST syntax trees of TypeScript code to extract API route signatures, type definitions, and generates OpenAPI 3.0 YAML specs.

#5 Git Repository Refactoring Pipeline
Difficulty: Hard

Create an automated code refactoring script that scans a codebase for deprecated API usages, rewrites imports using AST transformations, and formats git diff reports.

#6 LLM Function Calling Tool Executor Engine
Difficulty: Hard

Implement a type-safe tool execution engine that validates JSON schema tool arguments, enforces execution timeout limits, sanitizes outputs, and handles multi-turn tool loops.

#7 Distributed Web Crawler & Structure Extractor
Difficulty: Medium

Write an asynchronous agent crawler with rate-limiting queues, robots.txt compliance, URL deduplication, and automated JSON metadata extraction pipelines.

#8 Automated Incident Triage & Log Analyzer
Difficulty: Medium

Create an agentic log analyzer script that parses multi-server error tracebacks, clusters root causes using regex clustering, and auto-generates Slack markdown incident summaries.

#9 Database Query Optimizer & Index Recommender
Difficulty: Hard

Implement an agent script that analyzes SQL EXPLAIN query execution plans, identifies missing indexes, detects N+1 query antipatterns, and outputs optimized SQL rewrites.

#10 CI/CD Pipeline Security Policy Checker
Difficulty: Medium

Write a static analysis tool that scans GitHub Actions workflow files for untrusted script injection vulnerabilities, hardcoded secrets, and generates security remediation scripts.