Own the stability of services and infrastructure within your domain. You'll be the one monitoring, responding to incidents, and making independent decisions under pressure — not waiting to be told what to do. Beyond keeping things running, you're expected to actively reduce manual toil through scripting and AI tooling.Mô tả công việcOperations
Monitor and handle production incidents involving services, queues, and databases.
Triage alerts, classify severity, and decide when to escalate independently.
Use Grafana, Splunk, or equivalent tools to catch issues before they escalate.
Deploy
Execute deploys, rollbacks, and change requests across DEV, UAT, Staging, and Production.
Validate post-release and act quickly if something goes wrong.
Improvement
Build scripts and tools to automate repetitive operational tasks.
Apply AI tools (Claude Code or equivalent) to reduce manual effort in daily work.Yêu cầu công việcRequirements
2–4 years of hands-on experience in system operations or IT operations.Kubernetes: Real operational experience required.Kafka / RabbitMQ: Understand the mechanics and have solved real incidents.Monitoring: Proficient with Grafana, Splunk, or equivalent tools.Scripting AI: Solid foundation in Python, Bash, or Go; actively uses AI tools to automate and move faster.Database: Comfortable with basic queries and understands operational fundamentals.
Skills Attributes
Strong incident response skills — able to work through ambiguous, high-pressure situations without a playbook.
Makes independent decisions: classify severity, determine escalation path, act.
Ownership mindset — you follow through, not just flag and wait.
Proactive about reducing toil; automation is a reflex, not an afterthought.
No requirement for relevant working experience