05 set
|
Obsidian
|
Milano
ph3About the work /h3 pWe're building a high-quality dataset of human preference judgments on AI-generated frontend code.
Each task hands you a reference web page — crawled from the real internet, delivered as a full zipped site tree plus screenshots of its default view and, on some pages, additional states reached by hovering, clicking, or scrolling.
Alongside it come two model attempts, A and B, each a zipped self-contained site tree.
The models only ever saw the screenshots; they never had the source.
/p pYou download all three, run them locally, view each at a ****×**** viewport, interact with them to reach every required state, then open the source of both attempts and grade them against each other — on visual fidelity per state, and on how the code is actually constructed.
Structure and responsiveness are explicitly part of the rubric, not just the render.
/p pThis is evaluation work, not authoring.
The defining skill is not that you can build a page — it's that you can open someone else's page and tell how it was built and where it cheats.
/p h3Please read before applying /h3 ul lipEach unit takes roughly 2–3 hours and is timed.
This is not microtask work; if you can only offer scattered 15-minute windows, you will not be able to finish a unit.
/p /li lipYou need a real local development environment.
A tablet, a Chromebook, or a locked-down work machine that cannot run a local static server will not work for this project.
/p /li lipYou need at least one completed Mercor engagement, delivered in full.
We are not onboarding net-new experts to this project.
/p /li /ul h3What you'll do /h3 ul lipRender a reference page and two candidate replications at ****×**** and judge which is the closer reproduction, state by state.
/p /li lipDiff visual fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow.
/p /li lipRead the source of both attempts and grade construction quality — distinguishing a replication that is genuinely correct from one that merely looks correct at one viewport.
Hardcoded pixel offsets,
absolute positioning standing in for real layout, inline style soup, a single undifferentiated div tree, or a screenshot pasted in as an instead of a rebuilt section.
/p /li lipTest responsiveness: a nav bar that looks right at ****px but collapses at ****px is a defect, and you should be able to say precisely why.
/p /li lipWrite a specific, evidence-cited justification for every preference.
We need "B nests the article body in a single absolutely-positioned div, so the text overlaps the footer below ****px, while A uses normal document flow" — not "A looks closer." /p /li lipUse the "this task is broken" escape hatch with judgment: distinguish a task that genuinely cannot be completed from the screenshots provided from one that is merely hard.
Over-flagging and under-flagging are both failure modes.
/p /li /ul h3You're a fit if you have /h3 ul lip3–8 years of professional web development experience, shipping web interfaces for a living, primarily in frontend or full-stack work.
/p /li lipFluency across web eras.
The reference pages are real crawled sites — one is a small charter-fishing business built in the table-and-image-map tradition, another a corporate press-release page with stacked navigation rows and social share widgets.
If your entire career happened inside a modern component framework and you have never authored raw CSS or seen a table used for layout, you will misjudge many of these pages.
/p /li lipCommand of hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and legacy float- and table-based layouts you can read and reason about.
You should be able to look at a rendered layout and predict what's holding it together before opening DevTools.
/p /li lipBrowser DevTools as muscle memory — setting an exact viewport,
walking the element tree, checking computed styles, watching what a hover handler mutates.
/p /li lipCommand-line comfort: unzipping an archive, standing up a static local server because the relative asset paths demand it, and untangling a broken image reference rather than giving up and grading from the screenshot.
/p /li lipEnough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily.
/p /li lipProfessional written English.
Every task ends in a free-text justification, and a rating without a specific rationale is worth very little.
/p /li /ul h3Equipment /h3 ul lipA desktop or laptop that displays a ****×**** viewport.
/p /li lipAdministrator rights on your own machine, so you can install and run a local server.
/p /li /ul h3Nice to have /h3 ul lipPrior RLHF, preference-labeling, model-evaluation, or structured code-review work — the strongest single signal.
Rubric-driven comparison at volume needs almost no ramp here.
/p /li lipPixel-perfect design-to-code experience: agency work, design systems, template production.
Anyone who has had a designer reject a build over four pixels has exactly the fidelity eye this needs.
/p /li lipAccessibility expertise (ARIA, semantic landmarks, heading hierarchy) — you'll notice immediately when an attempt renders a heading as a styled .
.
/p /li lipFamiliarity with how LLMs fail at code generation.
/p /li lipWeb scraping, archiving, or DOM-parsing background — comfort with messy crawled site trees.
/p /li lipMore than one completed Mercor project, and availability in contiguous multi-hour blocks.
/p /li /ul h3Note: /h3 pthis seat is for practicing web developers.
Backend-only, ML/data-science-only, mobile-native-only, and DevOps-only engineers do not have the UI instincts this requires, however strong they are otherwise.
Designers who do not code cannot grade the source axes at all.
Framework-only engineers who have never authored CSS outside a component library will struggle with the legacy reference pages.
/p /p #J-*****-Ljbffr
📌 Frontend Code Evaluation Specialist (Milano)
🏢 Obsidian
📍 Milano