We are looking for an experienced SRE/DevOps or Infrastructure Engineer to join our client’s team and work on enterprise storage and virtual infrastructure in Prague. You will combine storage engineering with Python automation, infrastructure as code, and monitoring to improve the reliability and efficiency of the platform. The role includes solving performance and latency issues, automating operational workflows, and building tools that reduce manual work. You will also be involved in incident response, troubleshooting, and implementing long-term improvements.
Informace o projektu
- Area: Storage and Virtualization
- 3 days on-site - Prague
- Project is for someone who has strong experience with enterprise infrastructure, Python, Linux, and tools such as Ansible or Terraform, this is an opportunity to work on challenging infrastructure problems with a real impact.
Co pozice obnáší
- Enterprise Storage and Virtual Infrastructure Integration (OpenStack, VMware)
- Solve customer performance and latency issues, provide workflow-based recommendations, and perform storage architect design and implementation, including capacity planning.
- Use software development methods to automate workflows and eliminate manual toil.
- Remotely deploy storage using scripts and internal tooling.
- Troubleshoot storage issues or script failures.
- Lead incident response and manage interrupts during critical incidents.
- Build automation and tools in Python and shell scripting
- Develop Python based tools and automation to support self service workflows
- Automate repetitive operational tasks and enhance deployment script efficiency.
- Integrate with internal and external APIs to orchestrate infrastructure workflows (compute, storage, network)
- Configuration Management, Infrastructure, and Monitoring as Code
- Use tools such as Ansible, Terraform, Puppet, or similar tools to manage infrastructure declaratively.
- Maintain reusable playbooks/modules and templates for common infrastructure patterns
- Enforce configuration standards, security baselines, and repeatable deployments across environments
- Monitoring, observability, and reliability
- Implement and improve monitoring, alerting, and dashboards for infrastructure health (e.g., Prometheus, Grafana, ELK/Nagios, or similar tools)
- Define and track key metrics (availability, latency, capacity, error rates), and drive improvements based on data
- Participate in incident response, perform root cause analysis, and implement long term fixes and runbooks
- Maintain the observability codebase and actively develop new monitors and alerts.
- Collaboration and support
- Provide guidance on best practices for using infrastructure platforms (VMs, containers, storage, networking)
- Participate in an on call rotation and planned maintenance windows as needed
Co potřebujeme
Praxe v oboru
5 let a více
Odborná zkušenost
- 5+ years of experience in Infrastructure/Storage Engineering or SRE/DevOps roles supporting enterprise storage organizations or teams. Knowledge of Everpure products is a plus.
- 3+ years of hands on Python development for:
- Automation scripts and tools
- REST API integrations
- Data collection, reporting, and operational tooling
- Experience with at least one configuration management or IaC tool (e.g., Ansible, Terraform, Puppet, Chef)
- Linux systems administration knowledge (e.g., Ubuntu, CentOS/RHEL)
- Understanding of networking fundamentals (TCP/IP, DNS, DHCP, VLANs, routing basics)
- Experience with office tools such as Jira, Slack, Google Workspace, and others.
- Knowledge of monitoring and alerting on key storage metrics.
- Strong problem solving skills, ownership mindset, and clear written and verbal communication
Technologie
- Ansible, Terraform, Puppet, Chef
- Linux (Ubuntu, CentOS/RHEL
- TCP/IP, DNS, DHCP, VLANs
- Jira, Slack, Google Workspace
- SAN/NAS
- Pure Storage products FA and FB
- CI/CD tooling and pipelines (Jenkins, GitHub Actions, ArgoCD, GitLab CI)
Jazykové požadavky
- Čeština, Angličtina
Možnost prodloužení
ANO
Co můžeme nabídnout
- The resources of a large company with the spirit of a startup
- Commitment to long-term collaboration
- A personalized approach and the chance to get to know each other
- Opportunities for professional growth through projects, training events, and meetups
- Discounted IT courses
- Loan of work equipment—laptop, cell phone
- Access to our Hub
- Pleasant, modern offices in Prague 4’s Kavčí Hory neighborhood
