포지션 상세
Lunit, a portmanteau of ‘Learning unit,’ is a medical AI software company devoted to providing AI-powered total cancer care.
Our AI solutions help discover cancer and predict cancer treatment outcomes, achieving timely and individually-tailored cancer treatment.
[About the Team]
• Lunit's Software Engineering (hereafter, "SE") department develops the Lunit INSIGHT product. Within the SE department, the Core Product Engineering team performs the Backend/Application development needed to productize INSIGHT AI models. We translate product requirements into development requirements and build Software as a Medical Device (SaMD) that complies with medical device guidelines.
• Our team's core goal is to ensure the INSIGHT product has optimized performance and reliability so it can operate across diverse countries and deployment environments.
[About the Position]
• The Site Reliability Engineer role sits between development and operations, improving reliability, observability, automation, and operational systems so the INSIGHT product can run stably.
• Rather than simply performing operational tasks, you will discover recurring problems and operational inefficiencies in production, analyze their root causes, and translate the findings into automation, improved observability, and improved operational processes.
• You will collaborate with Lunit International's SRE team, understand differing operational environments and processes, and connect and align the global operating system with the domestic product development environment.
• Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services.
• Design and enhance cloud infrastructure, deployment, monitoring, and operational automation systems.
•Analyze cloud resource usage and cost, and continuously drive cost optimization while considering stability and performance.
• Diagnose, recover from, and perform root-cause analysis on incidents, and prevent recurrence through monitoring, alerting, runbooks, and automation.
• Reliably manage and improve operational configurations and security elements such as per-environment/per-customer settings, certificates, secrets, and access permissions.
• Understand the product domain and collaborate with development teams to build reliability in from the design and development stages.
• Collaborate in English with Lunit International's SRE team, understanding and aligning differing operational environments and processes.
• Build an on-call and incident-response system for 24/7 service operations, and participate in the actual on-call rotation to handle production incidents. As the team grows, evolve and lead a sustainable on-call operating model.
[Tech Stack]
• Scripting/Language: Python, Bash
• Cloud/Infrastructure: Azure, Linux, Docker, Kubernetes, Terraform, Bicep
• CI/CD: GitHub Actions, Azure DevOps Pipelines
• Observability/Operations: Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow, Custom Dashboards
• Product Environment: Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQ
• Collaboration: Jira, Confluence, Slack, GitHub
• Experience directly designing and operating Azure-based production environments
• Hands-on experience with Linux, networking, and containers
• Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems
• Experience analyzing production incidents based on monitoring, and performing recovery and recurrence-prevention activities
• Experience collaborating with product and development teams to balance stability and development velocity
• Ability to discuss and coordinate technical decisions fluently in English with overseas engineers and colleagues from diverse roles
Our AI solutions help discover cancer and predict cancer treatment outcomes, achieving timely and individually-tailored cancer treatment.
[About the Team]
• Lunit's Software Engineering (hereafter, "SE") department develops the Lunit INSIGHT product. Within the SE department, the Core Product Engineering team performs the Backend/Application development needed to productize INSIGHT AI models. We translate product requirements into development requirements and build Software as a Medical Device (SaMD) that complies with medical device guidelines.
• Our team's core goal is to ensure the INSIGHT product has optimized performance and reliability so it can operate across diverse countries and deployment environments.
[About the Position]
• The Site Reliability Engineer role sits between development and operations, improving reliability, observability, automation, and operational systems so the INSIGHT product can run stably.
• Rather than simply performing operational tasks, you will discover recurring problems and operational inefficiencies in production, analyze their root causes, and translate the findings into automation, improved observability, and improved operational processes.
• You will collaborate with Lunit International's SRE team, understand differing operational environments and processes, and connect and align the global operating system with the domestic product development environment.
주요업무
You will design and improve the reliability, observability, deployment automation, and cloud infrastructure operations of the Lunit INSIGHT product to ensure it runs stably. This role is not limited to performing predefined operational tasks. You will identify gaps in the team's current SRE capabilities and operational systems, then define and drive the improvements needed.• Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services.
• Design and enhance cloud infrastructure, deployment, monitoring, and operational automation systems.
•Analyze cloud resource usage and cost, and continuously drive cost optimization while considering stability and performance.
• Diagnose, recover from, and perform root-cause analysis on incidents, and prevent recurrence through monitoring, alerting, runbooks, and automation.
• Reliably manage and improve operational configurations and security elements such as per-environment/per-customer settings, certificates, secrets, and access permissions.
• Understand the product domain and collaborate with development teams to build reliability in from the design and development stages.
• Collaborate in English with Lunit International's SRE team, understanding and aligning differing operational environments and processes.
• Build an on-call and incident-response system for 24/7 service operations, and participate in the actual on-call rotation to handle production incidents. As the team grows, evolve and lead a sustainable on-call operating model.
[Tech Stack]
• Scripting/Language: Python, Bash
• Cloud/Infrastructure: Azure, Linux, Docker, Kubernetes, Terraform, Bicep
• CI/CD: GitHub Actions, Azure DevOps Pipelines
• Observability/Operations: Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow, Custom Dashboards
• Product Environment: Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQ
• Collaboration: Jira, Confluence, Slack, GitHub
자격요건
• 5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operations• Experience directly designing and operating Azure-based production environments
• Hands-on experience with Linux, networking, and containers
• Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems
• Experience analyzing production incidents based on monitoring, and performing recovery and recurrence-prevention activities
• Experience collaborating with product and development teams to balance stability and development velocity
• Ability to discuss and coordinate technical decisions fluently in English with overseas engineers and colleagues from diverse roles










