Storage Site Reliability Engineer (GPFS / Storage Ops) | Ingénieur(e) SRE Stockage (GPFS / Opérations de stockage)
Manage and support large-scale enterprise storage environments, specifically GPFS and NAS technologies. Automate operational processes using Python to improve platform reliability and performance.
- Hybrid
- Montréal, QC
- Posted Sep 3, 2026
- Apply by Oct 3, 2026
- 1 position
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Job summary
We are looking for a Senior Storage Site Reliability Engineer (SRE) to join a high-performing infrastructure team responsible for designing, automating, and maintaining large-scale enterprise storage platforms. This role is ideal for someone with strong Linux administration, storage technologies, and automation expertise who enjoys solving complex reliability and performance challenges. Hybrid work - 3 days in office in Montreal What You'll Do ✅ Manage and support enterprise storage environments, including GPFS and NAS technologies ✅ Automate operational processes and infrastructure tasks using Python ✅ Monitor, troubleshoot, and improve storage platform reliability and performance ✅ Support Linux-based environments and distributed storage systems ✅ Collaborate with network, cloud, and infrastructure teams to resolve complex issues ✅ Drive continuous improvement through automation, tooling, and operational excellence What We're Looking For ✔ Strong experience with Linux Administration and/or NAS technologies (File Systems, SMB) ✔ Knowledge of Software Defined Storage and Cloud technologies ✔ Experience with storage and Unix protocols including NFS, SMB, Fibre Channel, and iSCSI ✔ Strong Python scripting skills for automation and tooling development ✔ Solid understanding of TCP/IP networking, DNS, CIFS, and related technologies ✔ Experience supporting large-scale infrastructure in production environments ✔ Strong troubleshooting and problem-solving skills Top Skills 🔹 GPFS (IBM Storage Scale) 🔹 Linux Administration 🔹 Python Automation 🔹 NAS Storage (NFS/SMB) 🔹 Software Defined Storage (SDS) 🔹 TCP/IP Networking 🔹 Site Reliability Engineering (SRE) Nous recherchons un(e) Ingénieur(e) principal(e) SRE Stockage pour rejoindre une équipe infrastructure performante responsable de la gestion, de l'automatisation et de la fiabilité de plateformes de stockage d'entreprise à grande échelle. Ce rôle convient parfaitement à une personne passionnée par Linux, les technologies de stockage et l'automatisation. En mode hybride - 3 jours en présentiel au bureau à Montréal Responsabilités ✅ Gérer et soutenir les environnements de stockage d'entreprise, incluant GPFS et les technologies NAS ✅ Automatiser les processus opérationnels et les tâches d'infrastructure à l'aide de Python ✅ Surveiller, diagnostiquer et optimiser la performance et la fiabilité des plateformes de stockage ✅ Assurer le support des environnements Linux et des systèmes de stockage distribués ✅ Collaborer avec les équipes réseau, infonuagique et infrastructure ✅ Contribuer à l'amélioration continue grâce à l'automatisation et aux meilleures pratiques SRE Compétences recherchées ✔ Excellente expérience en administration Linux et/ou technologies NAS (systèmes de fichiers, SMB) ✔ Connaissance du stockage défini par logiciel (SDS) et des technologies Cloud ✔ Maîtrise des protocoles de stockage et Unix : NFS, SMB, Fibre Channel, iSCSI ✔ Excellentes compétences en Python pour l'automatisation et le développement d'outils ✔ Bonne compréhension des technologies TCP/IP, DNS, CIFS et des réseaux d'entreprise ✔ Expérience dans des environnements critiques à grande échelle ✔ Fortes capacités d'analyse et de résolution de problèmes Compétences clés 🔹 GPFS (IBM Storage Scale) 🔹 Linux Administration 🔹 Python Automation 🔹 NAS Storage (NFS/SMB) 🔹 Software Defined Storage (SDS) 🔹 TCP/IP Networking 🔹 Site Reliability Engineering (SRE)
What you’ll do
Manage and support large-scale enterprise storage environments, specifically GPFS and NAS technologies. Automate operational processes using Python to improve platform reliability and performance.
Requirements
Requires strong experience in Linux administration, NAS technologies, and Python scripting for automation. Candidates should be proficient in storage protocols like NFS, SMB, Fibre Channel, and iSCSI.
Listed skills
- Continuous ImprovementPreferred
- SoftwarePreferred
- Problem solvingPreferred
- ReliabilityPreferred
- Performance SupportPreferred
- Operational ExcellencePreferred
- fiabilitéPreferred
- DNSPreferred
- DevelopmentPreferred
- TCP/IPPreferred
- LinuxPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Gpfs
- Linux Administration
- Python Automation
- Nas Storage
- Software Defined Storage
- Tcp/ip Networking
- Site Reliability Engineering
Job areas
- Technology
- Software
- Engineering
- Consulting
Additional details
- Minimum experience
- 5+ years
- Apply by
- Oct 3, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Office presence
- 3 days per week
- Seniority
- Associate
- Application method
- Direct apply is available