포티투닷(42dot)-Senior AI Data Pipeline Engineer
포티투닷(42dot)-Senior AI Data Pipeline Engineer
포티투닷(42dot)-Senior AI Data Pipeline Engineer
1/3
포티투닷(42dot)경기 성남시경력 3-12년

Senior AI Data Pipeline Engineer

포지션 상세

About the Team & Mission

42dot의 AI 데이터 파이프라인 엔지니어는 전 세계에서 수집되는 데이터를 처리하고 관리하는 글로벌 데이터 파이프라인을 설계하고 확장합니다. 페타바이트(PB)급 데이터를 대규모 GPU 인프라에 안정적으로 전달하여, 핵심적인 AI 워크로드를 가동하는 고처리량 시스템을 구축하고 운영하게 됩니다.

At 42dot, our AI Data Pipeline Engineer architect and scale global data pipelines that ingest and process data from worldwide sources. You will design and operate high-throughput systems to reliably deliver petabyte-scale data to our large-scale GPU infrastructure, powering mission-critical AI workloads.

주요업무

• 다양한 AI 및 머신러닝 프로젝트를 지원하기 위한 고성능·고확장성 데이터 파이프라인 설계 및 구축
• 글로벌 데이터 가용성 및 원활한 동기화를 위한 멀티 리전(Multi-region) 데이터 인프라 아키텍처 설계 및 구현
• 여러 AI 프로젝트를 동시 지원할 수 있도록 복잡한 브랜칭 및 로직 격리가 가능한 유연한 파이프라인 아키텍처 개발
• Databricks 및 Spark를 활용한 대규모 데이터 처리 워크로드 최적화(처리량 극대화 및 비용 최소화)
• Kubernetes 기반 컨테이너 데이터 환경 유지 보수 및 고도화로 데이터 워크로드의 안정적 실행 보장
• AI 리서처 및 플랫폼 팀과 협업하여 고품질 데이터를 학습 및 평가 파이프라인으로 효율적으로 공급

• Design and build high-performance, scalable data pipelines to support diverse AI and Machine Learning initiatives across the organization.
• Architect and implement multi-region data infrastructure to ensure global data availability and seamless synchronization.
• Develop flexible pipeline architectures that allow for complex branching and logic isolation to support multiple concurrent AI projects.
• Optimize large-scale data processing workloads using Databricks and Spark to maximize throughput and minimize processing costs.
• Maintain and evolve the containerized data environment on Kubernetes, ensuring robust and reliable execution of data workloads.
• Collaborate with AI researchers and platform teams to streamline the flow of high-quality data into training and evaluation pipelines.

자격요건

• 대규모 AI/ML 데이터셋을 위한 프로덕션급 데이터 파이프라인 구축 및 운영 경험
• Apache Spark 및 Databricks 생태계 등 분산 처리 프레임워크에 대한 높은 숙련도
• Apache Airflow 등 워크플로우 오케스트레이션 도구를 활용한 복잡한 의존성 관리 및 실무 경험
• Kubernetes 및 컨테이너 기술을 활용한 데이터 처리 컴포넌트 배포 및 확장 능력
• Apache Kafka 등 분산 메시징 시스템을 활용한 고처리량 데이터 수집 및 이벤트 기반 아키텍처 이해
• Python을 활용한 시스템 레벨 최적화 및 수준 높은 프로그래밍 역량
• 보안과 확장성을 고려한 클라우드 네이티브 서비스 및 인프라 구축 best practices에 대한 이해
• 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력
• 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량

• Extensive professional experience in building and operating production-grade data pipelines for massive-scale AI/ML datasets.
• Strong proficiency in distributed processing frameworks, particularly Apache Spark and the Databricks ecosystem.
• Deep hands-on experience with workflow orchestration tools like Apache Airflow for managing complex dependency graphs.
• Solid understanding of Kubernetes and containerization for deploying and scaling data processing components.
• Proficiency in distributed messaging systems such as Apache Kafka for high-throughput data ingestion and event-driven architectures.
• Expert-level programming skills in Python for system-level optimizations.
• Strong knowledge of cloud-native services and best practices for building secure and scalable data infrastructure.
• Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems.
• Strong communication skills to effectively collaborate with cross-functional teams and external partners.

기술 스택 • 툴

태그

마감일

상시채용

근무지역

경기 성남시 수정구 창업로40번길 20, A동 42dot
본 채용정보는 원티드랩의 동의없이 무단전재, 재배포, 재가공할 수 없으며, 구직활동 이외의 용도로 사용할 수 없습니다.
본 채용 정보는 에서 제공한 자료를 바탕으로 원티드랩에서 표현을 수정하고 이의 배열 및 구성을 편집하여 완성한 원티드랩의 저작자산이자 영업자산입니다. 본 정보 및 데이터베이스의 일부 내지는 전부에 대하여 원티드랩의 동의 없이 무단전재 또는 재배포, 재가공 및 크롤링할 수 없으며, 게재된 채용기업의 정보는 구직자의 구직활동 이외의 용도로 사용될 수 없습니다. 원티드랩은 에서 게재한 자료에 대한 오류나 그 밖에 원티드랩이 가공하지 않은 정보의 내용상 문제에 대하여 어떠한 보장도 하지 않으며, 사용자가 이를 신뢰하여 취한 조치에 대해 책임을 지지 않습니다.
<저작권자 (주)원티드랩. 무단전재-재배포금지>