MCP & Plug In Connectors Expert
Remote · Contract · Generalist · $50-190/hr
About the Opportunity
A leading AI research organization is seeking advanced LLM power users with strong experience using MCP and (more importantly) plugins/connectors for real-world personal life tasks.
This project focuses on evaluating how well AI systems handle personalized, multi-step life tasks that require context, judgment, planning, and use of connected tools such as Google Drive, Expedia, Notion, and other plugins/connectors.
This role is ideal for people who use AI heavily in their personal lives and can do a better job replicating what AI could do if it weren’t available
What You’ll Do
You will help evaluate AI systems on complex personal workflows, including tasks across:
Personal health
Travel
Activity planning, including food and dining
Services, such as home repair
Career search
Other life organization workflows
Responsibilities may include:
Creating realistic prompts for complex personal-life tasks
Executing tasks and actions while recording your screen (required)
Using your personal plugins/connectors while you complete actions
Writing clear explanations of AI successes and failures
Judging whether AI outputs are practical, personalized, and well-reasoned
Identifying where models miss context, overreach, fail to use tools correctly, or produce unrealistic results
Creating and applying detailed rubrics to assess model performance
Who We’re Looking For
Strong candidates will have:
US-based only
Strong MCP experience and plug in / connector usage
Experience using LLM plugins/connectors such as Google Drive, Expedia, Notion, and similar tools, multiple times a week
Heavy personal usage of LLM products
An active, rich LLM account with regular usage and approximately 6+ months of history
Willingness to sign a data-share consent form via DocuSign
Experience using AI for multi-step planning, research, decision-making, or personal workflows
Strong written judgment and attention to detail
Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic
Experience writing and evaluating against rubrics
Extensive rubric experience is especially valuable, including 100+ hours on prior rubric projects involving rubric design, evaluation, and quality assessment.
Ideal Candidate Profile
The strongest candidates are LLM power users who are already using plug-in tools in their personal lives for high-context tasks such as trip planning, health research, home services, food and dining decisions, career planning, personal organization, or similar workflows.
Candidates who want to be more competitive for this and future opportunities are encouraged to proactively spend time learning MCP and using LLM plugins/connectors before applying.
Why This Work Matters
LLMs are quickly becoming personal assistants for everyday decisions, but truly useful AI needs to do more than produce generic advice. It needs to understand context, preferences, constraints, tradeoffs, and what success looks like in real life.
Your evaluations will help improve how AI systems support people with practical, high-context tasks across food, health, travel, productivity, careers, and life organization. This work directly contributes to making AI assistants more personalized, trustworthy, and useful for real-world personal workflows.
Engagement Details
Expected commitment: 20+ hours/week
Ramp-up: 1–2 days required
Turnaround expectation: Ability to complete tasks within 24 hours
Equipment: Desktop or laptop required; Chromebooks are not supported
Experts added to the project will begin in a trial period to assess project fit, quality, and consistency before being considered for ongoing tasking.
Please note: This project is still in its early stages, so there may be an initial delay before tasking begins.
Listing sourced from Mercor. Annotation Academy is independent of these platforms and does not guarantee work or pay. See our disclosures.