Skip to content
All projects
022021

Large-scale data processing platform

A containerized annotation platform handling a million-word dataset, orchestrated across a distributed team.

The problem

A Japanese corporate client needed over 1 million words of text annotated for machine learning, but manual annotation workflows couldn't scale to that volume or stay consistent across a distributed international team.

What I built

Architected and deployed a containerized data annotation platform, orchestrating processing workflows in Kubernetes so work could scale horizontally against the dataset. Docker-based packaging kept the annotation toolchain identical for every team member, regardless of where they worked from.

Outcome

  • 1M+

    Japanese words in the annotated dataset

  • Kubernetes

    Horizontal scaling for distributed processing

  • Global team

    Coordinated across international collaborators

Architecture

The shape of the system after the work. Inspect any component to see the part it plays.

Large-scale data processing platformTap to inspect

Select a component to see its role in the system.