Japanese Pdf Annotation Specialist (Milano)

Japanese Pdf Annotation Specialist (Milano)

22 set
|
Mercor
|
Milano

22 set

Mercor

Milano

ph3Fluent Language Skills Required: /h3 /brpJapanese.
Native fluency in Japanese, including full command of kanji, hiragana and katakana, is required for this position.
All annotation and transcription work is performed in Japanese.
/p /brh3Why This Role Exists /h3 /brpDocument understanding breaks down fastest in the languages that parsing and vision-language models rarely see.
This project builds training data for exactly those languages: Japanese, alongside Korean and five Indic scripts.
Each task takes a real, publicly available PDF page and produces a complete structural map of that page, paired with a faithful transcription of every text region in the original script.
/p /brpThe dataset deliberately concentrates on the material models handle worst: handwriting, dense multi-column layouts, vertical text, tables, diagrams, and mixed-script pages.
Documents are drawn from newspapers, textbooks, examinations, and everyday formats such as flyers, forms, manuals, menus, brochures, notices and worksheets, so that the corpus reflects the real diversity of Japanese documents rather than a narrow band of easily parsed ones.
/p /brpDelivered work is human-authored throughout.
Component identification, component typing, reading order and all transcription are performed by people, not generated by parsing models.
/p /brh3What You'll Do /h3 /brul /brlipOpen and check a task: pages are provided, so you do not source documents yourself.
We find the PDFs and upload them for you.
Before annotating, confirm the page is in Japanese, is legible, has real content, and shows no personal details /p /li /brlipAnnotate structure: identify and bound every meaningful region of the page - document title,



section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field - and assign each a component type and a reading-order index /p /li /brlipRecord relationships: link each region to the figure or table it belongs to through a parent component identifier /p /li /brlipTranscribe faithfully: reproduce all text exactly as it appears, including kanji, hiragana, katakana, furigana and handwritten content, flagging any region where the source is not legible /p /li /brlipCapture page metadata: language, document type, source, page dimensions, and flags for tables, formulas and handwriting /p /li /brlipReview a colleague's work: every task is reviewed end to end by a second Japanese expert, and experienced annotators take on that review /p /li /br /ul /brh3Who You Are /h3 /brul /brlipYou are a native Japanese speaker with full command of kanji, hiragana and katakana, including furigana and variant character forms /p /li /brlipYou have professional experience in interpretation, journalism, transcription, translation, editorial work, or comparable document-intensive work /p /li /brlipYou are exact: character-level accuracy matters more here than speed, and a single wrong character is a defect /p /li /brlipYou are systematic:



you apply a taxonomy consistently across hundreds of pages rather than improvising per document /p /li /brlipYou are comfortable with unfamiliar layouts: vertical text, multi-column newspapers, exam papers, handwritten forms /p /li /br /ul /brh3Nice-to-Have Specialties /h3 /brul /brlipAI training data: annotation, labeling, grading, or bilingual evaluation for training datasets /p /li /brlipTranscription and localization: MTPE, subtitling, bilingual QA, OCR correction or post-editing /p /li /brlipDocument production: typesetting, copy-editing, proofreading, or digitization of Japanese-language material /p /li /brlipScript and encoding: Unicode normalization, Japanese input methods, full-width and half-width forms, and kanji variant handling /p /li /br /ul /brh3What Success Looks Like /h3 /brul /brlipEvery meaningful region on the page is captured, correctly bounded and correctly typed /p /li /brlipReading order reflects how the page is actually read, including vertical text and multi-column layouts /p /li /brlipTranscriptions match the source character for character, in Japanese script rather than romaji /p /li /brlipYour tasks pass second-expert review the first time /p /li /brlipUnsuitable pages are flagged up front rather than after thirty minutes of work /p /li /br /ul /brh3Why Join Mercor /h3 /brul /brlipBuild the training data that makes document AI work in scripts it currently handles badly /p /li /brlipWork from real published Japanese documents rather than synthetic or templated pages /p /li /brlipQuality leads on this project: accuracy is the first measure, with handling time tracked alongside it /p /li /br /ul /p #J-*****-Ljbffr

📌 Japanese Pdf Annotation Specialist (Milano)
🏢 Mercor
📍 Milano

Candidati a questo annuncio

Mostra le tue capacità professionali all'azienda, compila il form e lascia un tocco personale nella lettera di presentazione, aiuterà il recruiter nella scelta del candidato.

Iscriviti a questa job alert:

Ricevi via email le nuove offerte di lavoro per: japanese pdf annotation specialist (milano) / milano

Iscriviti a questa job alert:

Ricevi via email le nuove offerte di lavoro per: japanese pdf annotation specialist (milano) / milano