Search Jobs

1 - 20 of 3182 Site Reliability Engineer jobs

* No exact matches found. Showing closest results instead
Sort by:
GR

Site Reliability Engineer

Groww

4-6 Years Bengaluru, Karnataka, India Full-time

Position: Site Reliability Engineer Location: Bengaluru About Groww At Groww, we re on a mission to make financial services simple, accessible, and transparent for every Indian. As one of India s fastest-growing financial platforms, we help millions take control of their financial future through a wide range of products. We re a team driven by ownership, radical customer-centricity, and a deep passion for challenging the status quo. From intuitive design to robust engineering, everything we build is grounded in what our customers need. If you re excited about building systems that power the future of finance in India, we d love to hear from you. Our Vision To empower every Indian with the knowledge, tools, and confidence to make sound financial decisions. Our goal is to be the most trusted financial partner for millions across the country. Our Core Values Customer Obsession We put our users first, always. Extreme Ownership We own everything we do, end-to-end. Simplicity We keep things simple, effective, and intuitive. Long-term Thinking We focus on sustainable, impactful decisions. Transparency We believe in open communication and collaboration. Role Overview: As a Site Reliability Engineer (SRE) at Groww, you will be responsible for ensuring our systems are highly available, performant, and secure. You will work closely with engineering and infrastructure teams to improve reliability, automate deployments, and manage mission-critical services that power our platform. Key Responsibilities: Monitor and troubleshoot issues related to system performance, availability, and security. Define and maintain SLIs, SLOs, and Error Budgets to improve system reliability. Use tools like Grafana to analyze and report on metrics and trace data. Participate in the on-call rotation for 24/7 support of production systems. Collaborate with developers to ensure scalability and reliability are built into new services. Roll out security and infrastructure features proactively. Manage automated deployments, version control, and release rollouts. Perform Root Cause Analysis (RCA) for incidents and implement long-term fixes. Optimize system performance, conduct capacity planning, and create recovery strategies. Identify and automate repetitive tasks to reduce toil. Leverage CI/CD tools such as Git, Jira, Jenkins to streamline development workflows. Requirements: 4 6 years of relevant experience in SRE, DevOps, or infrastructure engineering. Bachelor's or Master's degree in Computer Science or a related field. Strong background in Linux/Unix system administration and networking. Hands-on experience with cloud platforms like GCP or AWS. Proficiency in programming languages such as Python, Java, or Go. Experience with monitoring and alerting tools: Grafana, Prometheus, New Relic, etc. Familiarity with configuration management tools. Experience with Kubernetes, Docker, and container orchestration tools is a strong plus. Excellent problem-solving, communication, and team collaboration skills. Be a part of one of India s fastest-growing fintech startups. Build and scale systems that impact millions of users daily. Work with passionate, driven teammates who are redefining financial services. A culture that encourages continuous learning, ownership, and transparency. If you're ready to help shape the future of fintech infrastructure in India, Groww is the place for you. Let s build something extraordinary together. Qualification : Bachelor's or Master's degree in Computer Science or a related field

CI/CD AWS kubernetes Docker terraform
BY

Senior DevOps / Site Reliability Engineer

Blue Yonder

10-13 Years Bengaluru, Karnataka, India Full-time

Job Title: Senior DevOps / Site Reliability Engineer Location: Pune, India Company: Blue Yonder Experience: 10 to 13 years Education: Bachelor s Degree in Computer Science, Engineering, or related STEM fields Company Overview Blue Yonder is a leading AI-driven Global Supply Chain Solutions provider and consistently recognized as one of Glassdoor s Best Places to Work. We are driving the next wave of digital transformation in manufacturing and retail, delivering innovative SaaS solutions that power intelligent supply chains across the globe. We are looking for a Senior DevOps / Site Reliability Engineer (SRE) to lead the design, development, deployment, and operational management of our Azure SaaS solution. This role requires strong DevOps, cloud delivery, and infrastructure automation expertise, along with leadership capabilities to guide a growing global team. Role Overview In this role, you will be responsible for architecting, planning, and executing end-to-end delivery pipelines, supporting both product development and operational stability. Working closely with platform, product, and architecture teams, you will implement best-in-class DevOps and SRE practices, ensuring scalability, resilience, and cost optimization. Key Responsibilities Architect, design, and manage CI/CD pipelines and infrastructure for a cloud-native, multi-tenant SaaS solution on Azure. Lead sprint planning, backlog grooming, and architecture discussions. Develop quality automation scripts and tools to reduce manual efforts and enable self-healing, self-service capabilities. Identify and resolve operational bottlenecks and proactively improve observability (monitoring, alerting, logging). Participate in code reviews, ensure secure and scalable designs, and mentor junior and mid-level engineers. Collaborate with stakeholders to understand business and technical requirements and translate them into actionable user stories. Implement and enforce cloud cost optimization strategies. Conduct post-incident reviews with a blameless culture to identify root causes and drive continuous improvements. Automate service requests and standard operational procedures. Drive improvements to the team s continuous integration pipeline, ensuring rapid and reliable deployments. Stay updated with the latest DevOps, SRE, and cloud technologies and bring innovative ideas to the table. Participate in team hiring and actively contribute to onboarding new team members. Technical Environment Languages: Java, Python, PowerShell, Shell Scripting DevOps Tools: Azure DevOps, GitHub Actions, Jenkins Cloud: Microsoft Azure (ARM Templates, AKS, Event Hub, HDInsight, Azure AD, Application Gateway, Virtual Networks) Architecture: Microservices, Kubernetes, Docker, Event-driven architecture Frameworks: Spring Boot, Hibernate Monitoring & Logging: Elasticsearch, Spark, Kafka Databases: RDBMS, NoSQL Version Control: Git Required Skills & Experience Bachelor s Degree (STEM preferred) with 10 to 13 years of experience in DevOps, Cloud Delivery, or Site Reliability Engineering. Proven hands-on experience with Azure Cloud Services. Expertise in setting up and optimizing CI/CD pipelines. Strong scripting experience: Shell and PowerShell are mandatory; Python is a plus. Strong understanding of container technologies (Docker, Kubernetes) and microservices architecture. Experience integrating and managing third-party monitoring and logging tools. Strong problem-solving skills and ability to work with global, cross-functional teams. Excellent communication and stakeholder management skills. Nice to Have Development experience in Java or Python. Experience working in agile teams with a product-centric mindset. Experience working in manufacturing or retail domains. Exposure to AI/ML-driven monitoring and observability tools. Work with cutting-edge technologies on globally impactful solutions. Collaborate with diverse and talented teams across the US, India, and the UK. Foster your career growth through mentorship, continuous learning, and leadership opportunities. Experience an inclusive, flexible work culture where innovation and creativity thrive. Diversity, Inclusion, Value & Equality (DIVE) At Blue Yonder, we are committed to building an inclusive environment where everyone feels empowered to be themselves. All qualified applicants will receive consideration for employment regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status. Qualification : Bachelors Degree in Computer Science, Engineering, or related STEM fields

CI/CD AWS kubernetes Docker terraform
CO

Senior Site Reliability Engineer

Couchbase

5+ Years Bengaluru, Karnataka, India Full-time

Job Title: Site Reliability Engineer (SRE) Cloud Platform & Production Pipeline Initiatives Location: Bangalore, India (Office-based role) About Couchbase: As industries race to embrace AI, traditional database solutions fall short of rising demands for versatility, performance, and affordability. Couchbase is leading the way with Capella, the developer data platform for critical applications in our AI-driven world. By uniting transactional, analytical, mobile, and AI workloads into a seamless, fully managed solution, Couchbase empowers developers and enterprises to build and scale applications with unmatched flexibility, performance, and cost-efficiency from cloud to edge. Trusted by over 30% of the Fortune 100, Couchbase is unlocking innovation, accelerating AI transformation, and redefining customer experiences. Come join our mission! Job Overview: As a Site Reliability Engineer (SRE), you will play a pivotal role in managing, optimizing, and maintaining Couchbase s cloud infrastructure for Capella, our Database as a Service (DBaaS) platform. You will be responsible for ensuring the reliability and performance of our cloud service while collaborating closely with engineering teams to improve deployment pipelines, security practices, and overall system health. You will work across cloud platforms and multiple tools to provide guidance, mentorship, and contribute to the strategic direction of cloud operations. Responsibilities: Infrastructure Management: Manage, monitor, and maintain the infrastructure for Capella to ensure reliable operations. Security & Compliance: Implement and manage cloud environments in accordance with company security guidelines, including vulnerability management, penetration testing, and compliance requirements (SOC 2, PCI-DSS, GDPR, HIPAA, etc.). CI/CD & Release Pipeline: Collaborate with engineering teams to optimize CI/CD processes, aiming for a highly resilient deployment strategy, ideally with zero downtime. Cloud Optimization: Stay up-to-date with new technologies and industry trends to continuously improve cloud platform architecture and meet the evolving needs of the business. Security Integration: Work with development teams to integrate security scanners within the DevOps lifecycle, enhancing security posture. Leadership & Mentorship: Provide guidance on architecture, code reviews, and technical feedback to improve service reliability, security, cost, and performance. Incident Management: Demonstrate exceptional problem-solving skills, proactively identifying and addressing potential issues before they affect business operations. Collaboration: Partner with development teams, application owners, and stakeholders to integrate best practices and ensure seamless service delivery. Requirements: Experience: 5+ years in Site Reliability Engineering (SRE), DevSecOps, or similar roles, with significant experience working in public cloud environments. Programming & Scripting: Proficiency in languages such as Go, Python, Java, or Ruby. Linux Expertise: High proficiency with Linux operating systems. Kubernetes Management: Experience in managing and maintaining Kubernetes clusters (both self-managed and managed platforms like AWS EKS). Security & Vulnerability Management: In-depth knowledge of security tools and practices (vulnerability management, pen testing, SCA, DAST, SAST), with hands-on experience using tools like Sysdig, Synk, and Blackduck. Cloud Platforms & Tools: Strong experience with cloud platforms (AWS, GCP, Azure) and open-source tools like Artifactory, Jira, Jenkins, Grafana, Prometheus, Datadog, Thanos, etc. Configuration Management: Proficiency with Terraform, Git, and CI/CD platforms (e.g., CircleCI, GitHub, Spinnaker). Networking Security: Solid understanding of TCP/IP, DNS, HTTP, Firewalls, VPNs, and other networking security concepts. Preferred Skills: Availability & Reliability: Knowledge of SLO/SLA, availability, reliability, and performance concepts. Incident Management: Experience with on-call rotations and incident management. Database Experience: Familiarity with databases, particularly Couchbase. Security Certifications: Relevant certifications in security or cloud technologies are a plus. Couchbase reimagines database technology to deliver a fast, flexible, and affordable cloud database platform, empowering developers to build applications with exceptional customer experiences. Trusted by over 30% of the Fortune 100, Couchbase drives innovation and customer success through its Capella platform. Benefits at Couchbase: Generous Time Off Program: Flexibility to care for yourself and your family. Wellness Benefits: Access to world-class medical plans, dental, vision, life insurance, and employee assistance programs. Financial Planning: RSU equity program, ESPP, retirement planning, and business travel insurance. Career Growth: Focused on your career development and success. Fun Perks: Ergonomic and comfortable office setup, food & snacks for in-office employees, and more!

AWS kubernetes Docker CI/CD terraform
NV

Senior Site Reliability Engineer

Nvidia

7+ Years Pune, Maharashtra, India Full-time

NVIDIA s Infrastructure, Planning and Processes (IPP) organization is seeking a hard-working and experienced Site Reliability/DevOps Engineer, with strong background in Infrastructure Management, Monitoring, Automation, & System Administration, to join our Sanity Operations Team in Pune. The IPP Org provides Infrastructure, Products & Services for multiple software teams including GPU, Mobile, and Automotive divisions working on Nvidia's extraordinary products & services. The team is responsible for hosting, enabling & running the large scale private cloud systems & services, for our in-house Testing CI framework. The cloud hosts a heterogeneous mix of machines and devices with various operating systems (Windows/Linux/Android, etc.), running with NVIDIA GPUs and Tegra Processors. What you ll be doing: Create resilient, scalable, and efficient test and deployment pipelines. Design and implement complex automation platforms to identify & resolve operational inefficiencies. Triaging software, hardware and infrastructure issues and maintaining high availability for our infrastructure & services. Deploying & Monitoring critical high performance, large scale services running on Geo-distributed systems. Continuously Strive for efficient utilization & management of the infrastructure. Automate processes for enabling developers to adopt self-service practices, while ensuring compliance with security standards. Work with architects and engineers across the teams to review the designs & solutions during development and deployment phases. Collaborate with our other engineering teams to deliver reliable, robust, and high-performance capability of the underlying infra. Mine & analyze data from multiple sources for identifying scaling & optimization opportunities. What we need to see: Bachelor s or Master s degree in computer science, Software Engineering, or equivalent experience with 7+ years of experience in a DevOps environment. Strong hands-on experience in Configuring, maintaining, and building upon deployments of industry-standard tools (e.g. Kubernetes, Jenkins, Docker, CMake, Gitlab, Jira, etc) Working Experience in monitoring & maintaining large-scale infrastructure applications running in a microservice-based architecture. Proficient with Virtualization architecture with strong experience in Kubernetes, VMs, Dockers. Experience with continuous integration and continuous delivery systems such as GitLab, GitOps, Jenkins, Packer, and Terraform. Strong Python scripting skills, with proven background of using/writing JSON/REST APIs. Fluency in using MySQL or equivalent NoSQL databases queries Solid understanding of configuration management tools like, Chef, Puppet, Ansible, etc. Working Experience with Perforce, GIT or any other version control system is necessary. Experience with telemetry and alerting systems such as Kibana, Elastic Search, Grafana, and Prometheus to create rich visualizations of system health over time. Ability to self-manage, show leadership, mentor others and communicate well. Ways to stand out from the crowd: Understanding of networking concepts like TCP/IP and firewall management. Exposure to web apps/dashboards on frameworks like Django, AngularJS, VueJS, etc. High level understanding of Build and Test systems. Experience in Building regression detection systems by analyzing real-time production data, emphasizing important metrics. Innovating with industry-standard tools and collaborating with the open source community Qualification : Bachelors or Masters degree in computer science, Software Engineering, or equivalent experience with 7+ years of experience in a DevOps environment.

kubernetes Docker grafana
ES

Lead Site Reliability Engineer (azure)

Epam Systems

5+ Years Pune, Maharashtra, India Full-time

We are seeking a highly skilled and experienced Lead Site Reliability Engineer with a focus on Azure environments to join our team. In this crucial role, you will leverage your expertise to enhance the reliability and scalability of our cloud-based platforms, ensuring efficient operation and optimal performance. This position involves collaborating closely with cross-functional teams to migrate existing services to the OpenShift platform and make our infrastructure Cloud agnostic. As a leader, you ll guide your team in creating resilient systems and processes that support both internal and external customers relying on our desktop applications and services. Responsibilities Oversee migration of services to OpenShift and work towards making our infrastructure Cloud agnostic Run pipelines using Azure DevOps for environment configuration and application deployment Leverage Python, bash, and PowerShell to automate routine and complex tasks Implement and manage Kubernetes and container-based environments Monitor cloud resources efficiently and improve system performance in line with SLI metrics Debug and resolve operational issues swiftly and effectively Collaborate with development and operations teams to ensure system reliability and security Mentor team members and lead by example in maintaining best practices for site reliability Continuously assess, improve and optimize existing system architecture and applications Stay up-to-date with technological advancements and integrate innovative tools and techniques Requirements 5+ years of experience as a Systems Engineer with a development background 1+ years of relevant leadership experience Proficiency in Linux and Docker with hands-on experience in Kubernetes Capability to use at least one of the following scripting languages: Python, Bash, PowerShell Background in infrastructure management including networking and operating systems Familiarity with monitoring tools in cloud environments and understanding of SLI concepts Familiarity with Azure and/or GCP as cloud service providers Nice to have Experience working with Windows Knowledge of CI/CD pipelines, particularly Azure DevOps Understanding of Istio and GitOps tools like ArgoCD We offer Opportunity to work on technical challenges that may impact across geographies Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications Opportunity to share your ideas on international platforms Sponsored Tech Talks & Hackathons Unlimited access to LinkedIn learning solutions Possibility to relocate to any EPAM office for short and long-term projects Focused individual development Benefit package: Health benefits Retirement benefits Paid time off Flexible benefits Forums to explore beyond work passion (CSR, photography, painting, sports, etc.) Qualification : 5+ years of experience as a Systems Engineer with a development background

kubernetes Docker
II

Site Reliability Engineer - Z Platform

Ibm India

Fresher Bengaluru, Karnataka, India Full-time

Introduction: The IBM CIO Technology Platform Transformation team plays a crucial role in modernizing IBM's technology infrastructure and platforms. By leveraging emerging technologies such as AI, machine learning, and cloud computing, the team aims to enhance security, streamline processes, and improve user experience. The team's mission is to optimize IT functions, reduce technical debt, and drive automation while fostering a culture of innovation and continuous improvement. In this role, you will be joining the CIO Hybrid Cloud Z Platform & Strategy Team, where you will maintain, support, and enhance multiple aspects of the Z environment, including z/OS storage management, performance, and networking tools. You will play a key role in automation, identifying improvements, and collaborating with various teams to improve operational efficiency. Responsibilities: Technical Support & Problem Resolution: Provide problem determination and source identification to resolve technical issues within the Z environment. Automation & Optimization: Recommend and implement optimization strategies and automation processes for technical support tools, procedures, and systems. Collaboration: Work closely with global teams to diagnose, prioritize, and resolve issues, providing guidance and fostering collaboration across various teams. Mentorship & Training: Offer technical training and mentorship to other members of the team, sharing knowledge and best practices. System Design & Implementation: Lead system design discussions for z/OS systems, plan and implement new solutions, and ensure they align with business goals and objectives. Continuous Improvement: Identify points of improvement in technical processes and propose innovative solutions through automation to enhance operational efficiency. Required Education & Experience: Bachelor's Degree in a relevant field (required). Master's Degree (preferred). Technical Expertise: Deep Knowledge of z/OS & z/OS Storage Support: Proficient in managing z/OS and storage in a mainframe environment, including installation, configuration, high availability, performance tuning, and security. Cloud Infrastructure & Network Knowledge: Experience in cloud infrastructure and network technologies, particularly in the context of z/OS environments. Problem Solving & Autonomy: Strong problem-solving skills, with the ability to work autonomously, meet goals, and apply innovative thinking to solve complex problems. Global Team Collaboration: Experience working with global teams across different locations, contributing to a collaborative and solution-driven environment. System Design & Leadership: Proven ability to lead system design discussions and plan and implement z/OS system configurations and changes. Fluent in English: Strong written and verbal communication skills in English. Preferred Technical Experience: REXX Programming: Experience with REXX programming for automation and scripting in a z/OS environment. Ansible on Mainframe: Familiarity with Ansible for automation on Mainframe systems. Zowe, ZOAU, R3S Knowledge: Knowledge of Zowe, ZOAU, and R3S for modernizing mainframe environments. Middleware & Systems Management Experience: Experience with z/OS middleware systems such as DB2, IMS, Base/Storage/TWS/Network/Netview. Automation & Testing: Experience in automating workloads, performing test automation, and optimizing operational workflows. Basic Container Technology Knowledge: Familiarity with container technologies and tools such as Docker, Kubernetes, and Red Hat OpenShift. About the Business Unit: The IBM Finance Organization is responsible for driving enterprise performance and transformation. As the financial stewards of IBM, we deliver IBM's financial strategy, develop new business models, and mitigate enterprise risk. The group is focused on creating value and improving the financial aspects of IBM s business across a variety of sectors, including accounting, financial planning, business controls, tax, treasury, and business development. Why Join IBM? Innovative Environment: Work on cutting-edge technologies and collaborate with experts in the field. Global Impact: Contribute to the transformation of IBM's hybrid cloud platform and technology infrastructure. Career Development: Gain exposure to advanced systems and automation techniques while working with global teams. Competitive Compensation & Benefits: Enjoy competitive pay, performance-based rewards, and a range of employee benefits. If you're passionate about technology, innovation, and driving automation in a hybrid cloud environment, join the IBM CIO Hybrid Cloud Z Platform & Strategy Team today!

System Monitoring kubernetes Docker prometheus
I(

Site Reliability Engineer -- Logging And Monitoring

Ibm (international Business Machines)

2-5 Years Bengaluru, Karnataka, India Full-time

Introduction A career in IBM Software means you ll be part of a team that transforms our customer s challenges into solutions. Seeking new possibilities and always staying curious, we are a team dedicated to creating the world s leading AI-powered, cloud-native software solutions for our customers. Our renowned legacy creates endless global opportunities for our IBMers, so the door is always open for those who want to grow their career. IBM s product and technology landscape includes Research, Software, and Infrastructure. Entering this domain positions you at the heart of IBM, where growth and innovation thrive. Your role and responsibilities In this role, you will build and maintain an observability stack for IBM s Cloud Object Storage service using managed services as well as custom built services. This stack is used by Cloud Object Storage SREs and devs to understand the health of the service. Work duties and responsibilities include: Design, setup, configure and implement the COS Monitoring System using technologies such as Elasticsearch, Logstash, Kibana, Kafka, Kafka Mirrors, Filebeat, Grafana and Sysdig. Automate CICD tasks and infrastructure using Ansible, Terraform, Jenkins, and Travis. Experience with microservices and distributed application architecture, such as containers and Kubernetes. Experience with Linux administration and programming languages such as java, python and sql. Performance and configuration tuning to support the increasing load of data flowing into the COS Monitoring System. Provide design recommendations and thought leadership to provide best-in-class observability as part the COS Monitoring System. Provide 24x7 on-call customer support on a rotational basis. Design and develop dashboards for metrics analysis Design, Develop and Configure an alerting solution for an end-to-end incident management and recovery process by integrating Sysdig with Pagerduty, Email and Slack. Required education Bachelor's Degree Preferred education Bachelor's Degree Required technical and professional expertise Ability and tenacity to solve increasingly complex technical issues through analysis and a variety of problem-solving techniques. Working knowledge of Object-Oriented Python with demonstrable experience in applying these skills. Working knowledge of Linux environments. Experience working in an Agile-Scrum development environment. Experience using tools such as Jira, GitHub and Logging and monitoring tools BS in CS, CE or similar field, plus 2 to 5 years relevant work experience. Qualification : BS in CS, CE or similar field, plus 2 to 5 years relevant work experience.

grafana kubernetes Docker
CE

Sr Devops Engineer

Celigo

5-10 Years Hyderabad, Telangana, India Full-time

Job Title: DevOps Engineer Location: Hyderabad, India About Celigo: At Celigo, we are pioneering the future of application integration with innovative strategies, cutting-edge technologies, and a dedicated team. Our core mission is simple: to enable independent, best-of-breed applications to work together seamlessly. We believe that every department and business end user should have choices in software selection, and integration challenges should never be an obstacle. Your Role: Celigo powers integration for a large number of customers worldwide through our Integrator.io iPaaS platform, processing millions of transactions daily. Our services are mission-critical, large-scale solutions in the area of iPaaS. We are looking for passionate DevOps engineers who thrive on managing modern, large-scale production services on the cloud. The successful candidate will work in alignment with Celigo's DevOps charter to deliver high-quality products faster, safer, and in compliance with company norms using automation, tooling, and processes. Education: Master's/Bachelor's degree in Computer Science/Engineering, Software Engineering, or a related discipline. Experience: 5-10 years of total experience in software product development organizations, with at least 3 years of experience in DevOps. Role Expertise: Proven experience as a Site Reliability Engineer (SRE) or similar role, including hands-on experience in managing mission-critical, large-scale product operations such as provisioning, deployment, upgrades, patching, and incidents in production on cloud. Cloud Skills: Must have working knowledge of AWS services (e.g., VPC, EC2, EKS, S3, IAM). Automation Expertise: Proficient in Infrastructure as Code (IaC) using Terraform and scripting with Python/Bash. Experience with version control systems like Git. Configuration Management: Familiarity with automation tools like Chef, Ansible, and Puppet. Security: Basic understanding of security compliance standards and regulations such as SOC2, HIPAA, GDPR. CI/CD Expertise: Hands-on experience in designing, delivering, and maintaining CI/CD pipelines using tools like Jenkins, Travis CI, Spinnaker, and ArgoCD. Monitoring & Telemetry: Experience with observability tools like Splunk, ELK, and Kafka. Database Knowledge: Familiarity with MongoDB and Kafka. Problem Solving & Troubleshooting: Strong problem-solving, troubleshooting, and analytical skills, demonstrated in previous roles. Communication Skills: Excellent verbal and written communication skills. Agile Experience: Experience working in an Agile development environment. Enterprise Monitoring Systems: Experience with enterprise-level monitoring systems is highly desirable. Platform Deployment: Build and deploy the Integrator.io platform on production, staging, and other environments as required on the cloud. Automation: Use automation tools to eliminate manual steps, including infrastructure provisioning, code deployment, and monitoring/notifications. Tool Administration: Research, deploy, and manage tools such as Splunk and Kafka to support the Integrator.io platform. Collaboration with Security & Compliance Teams: Work closely with the security team to gather data points for audits and address any identified security gaps through programs like HackerOne. CI/CD Pipeline: Design and build CI/CD pipelines for Integrator.io, collaborating with senior architects and engineers. Engineering Excellence: Continuously raise the bar of engineering excellence by implementing best practices in DevOps. The Best Candidate: Passionate about building a world-class DevOps organization. Demonstrates the ability to manage large distributed systems in production on the cloud. Understands how infrastructure components work together. Has experience working with globally distributed teams. Enjoys solving complex problems and is excited about learning new techniques in cloud-based distributed systems. Able to automate tasks using high-level programming languages. Thrives in a fast-paced, dynamic environment and is adaptable to shifting priorities. Quick to learn, and knows when to listen and when to take charge. Qualification : Masters/Bachelors degree required in Computer Science/Engineering, Software Engineering or Equivalent discipline.

AWS Docker kubernetes Python jenkins
OR

Site Reliability Developer 2/3

Oracle

5+ Years Bengaluru, Karnataka, India Full-time

Job Description: Site Reliability Engineer - OCI Cloud Engineering Team Role: Site Reliability Engineer (SRE) Team: OCI OLTP (Online Transaction Processing) Location: Kiev Career Level: IC2 Experience: 5+ years Overview: Oracle Cloud Infrastructure s (OCI) OLTP organization is seeking a Site Reliability Engineer (SRE) to join our dynamic and fast-paced Cloud engineering team. The team is responsible for mission-critical distributed systems and cloud services, and we are looking for an engineer who is deeply interested in databases, distributed systems, and cloud services. If you thrive in an environment where innovation, problem-solving, and operational excellence intersect, this is an exciting opportunity for you! As a member of the SRE services, you will focus on Cloud Services, building deployments, operations, security vulnerability mitigation, and automation. You will be instrumental in fostering a culture of Site Reliability Engineering (SRE) within the team, and your work will directly contribute to ensuring the stability, performance, and reliability of Oracle s global cloud service infrastructure. This role requires someone who is adaptable, highly motivated, and capable of managing large-scale cloud environments with a focus on continuous improvement. Key Responsibilities: Cloud Service Operations & Reliability: Deploy, operate, and maintain large-scale cloud service products in a highly available, fault-tolerant, and scalable environment. Collaborate with internal teams to identify and mitigate cross-team issues that pose operational risks to cloud services. Focus on systems reliability and ensure the continuous availability of cloud services by automating tasks and eliminating manual interventions. Automation & Improvements: Automate operational tasks and improve service deployments, focusing on scaling, performance, and uptime. Contribute to CI/CD systems, ensuring seamless integration and continuous delivery for cloud-based services. Leverage automation tools such as Terraform, Grafana, and Bitbucket to streamline operations. Security & Incident Response: Mitigate security vulnerabilities within cloud services and ensure compliance with Oracle's security standards. Participate in on-call rotations to provide immediate troubleshooting support and ensure rapid issue resolution. Perform deep analysis of service performance and collaborate with team members to diagnose and resolve issues that affect service availability or performance. Collaborative Problem-Solving: Work closely with cross-functional teams, including development, database, networking, and storage experts, to ensure the reliability and performance of services. Identify systemic issues and potential risks, develop solutions, and ensure proper documentation and communication with stakeholders. Documentation & Knowledge Sharing: Contribute to documentation such as runbooks, operational guides, and troubleshooting manuals. Mentor junior engineers and share knowledge on best practices for site reliability engineering and cloud service operations. Continuous Learning: Stay up to date with new cloud technologies, trends, and best practices, and actively implement them in your day-to-day work. Technical and Professional Requirements: Cloud Services & Infrastructure: 5+ years of experience in SRE, DevOps, or Automation roles with a focus on large-scale infrastructure and cloud services. Hands-on experience with cloud platforms (e.g., OCI, AWS, Azure) and expertise in compute, database, networking, and storage services within cloud environments. Automation & Tooling: Proficiency with automation tools such as Terraform, Grafana, LumberJack, and Shepherd. Solid experience in using CI/CD tools and processes for cloud service deployments and operations. Scripting & Systems: Strong knowledge of scripting languages, particularly Python and Java. Familiarity with Linux systems, docker containers, virtualized infrastructure, and orchestration (e.g., Kubernetes). Performance & Troubleshooting: Excellent troubleshooting skills with a focus on performance, availability, reliability, and scalability of distributed systems. Experience in operating fault-tolerant, highly available, high-throughput distributed systems. Security & Incident Management: Familiarity with security practices and mitigating security vulnerabilities in cloud services. Proven ability to handle incident response and provide efficient troubleshooting during on-call rotations. Collaboration & Communication: Strong verbal and written communication skills, capable of working effectively with diverse teams across multiple geographies. Ability to work in a highly collaborative environment, driving operational excellence and customer satisfaction. Preferred Qualifications: Experience in operating and maintaining multi-tenant, cloud-based infrastructure with a focus on scalability and high availability. Familiarity with tools and platforms like Grafana, Prometheus, and other observability and monitoring tools. Experience in networking and storage technologies in a cloud environment. Joining OCI s OLTP team as an SRE gives you the opportunity to work with cutting-edge technologies and contribute to the operational excellence of Oracle s global cloud infrastructure. This is a chance to grow your skills in a highly dynamic environment and to solve complex problems that directly impact mission-critical cloud services. With a focus on automation, scalability, and high performance, you will be an essential part of a team that powers Oracle s leading cloud services. If you are an experienced engineer passionate about cloud technologies, automation, and ensuring the reliability of large-scale systems, we encourage you to apply and join us in this exciting journey!

AWS kubernetes Docker prometheus grafana
QU

Autoit Solutioning Engineer, Lead

Qualcomm

Fresher Bengaluru, Karnataka, India Full-time

Job Title: Site Reliability Engineer (SRE) General Summary: We are seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our dynamic team. This role is critical in ensuring the stability, scalability, and security of our infrastructure and services. As an SRE, you will work collaboratively with software engineers, data scientists, and product managers to optimize system reliability while driving automation and continuous improvement. You will be responsible for modernizing traditional services, implementing cutting-edge technology, and proactively managing infrastructure to maintain operational excellence. If you are passionate about automation, DevSecOps, system performance, and infrastructure resilience, this role offers an exciting opportunity to make a meaningful impact. Key Responsibilities: System Monitoring & Incident Response: Continuously monitor system health, detect anomalies, and respond to incidents promptly. Investigate and troubleshoot service-related issues, ensuring minimal disruption. Implement proactive measures to prevent downtime and optimize system stability. Infrastructure Automation & DevOps Implementation: Develop and maintain Infrastructure-as-Code (IaC) scripts to automate deployments and scaling. Automate routine operational tasks to improve efficiency and reduce manual intervention. Leverage DevSecOps practices to ensure secure and resilient deployments. Performance Optimization & Capacity Planning: Collaborate with development teams to enhance software performance and system responsiveness. Identify and resolve system bottlenecks to improve speed, efficiency, and reliability. Forecast resource requirements based on traffic patterns and business growth. Security, Compliance & Risk Management: Implement security best practices and compliance measures across all infrastructure layers. Conduct security audits and ensure systems meet industry-standard security guidelines. Proactively assess and mitigate risks associated with infrastructure and deployments. Required Qualifications & Skills: Technical Expertise: Extensive experience with Linux-based environments (Ubuntu, RedHat), including system administration and troubleshooting. Strong proficiency in scripting and automation using Python, Bash, or Go. Experience with containerization and orchestration technologies such as Docker and Kubernetes. Familiarity with CI/CD pipelines and tools like Jenkins, Puppet, Vault, and Splunk. Hands-on experience with cloud platforms (AWS, Azure, or GCP). Problem-Solving & Leadership: Strong analytical skills with the ability to diagnose and resolve complex system issues. Self-driven, highly motivated, and able to work independently in a fast-paced environment. Ability to collaborate cross-functionally and communicate technical solutions effectively. Security & Reliability Focus: Solid understanding of DevSecOps principles and secure system design. Ability to implement monitoring, logging, and alerting solutions to maintain system resilience. Passion for continuous learning and leveraging data-driven approaches for system improvement. Work in a high-impact role that directly contributes to the reliability and scalability of mission-critical systems. Be part of an innovative, forward-thinking team that values automation, collaboration, and continuous improvement. Competitive salary, professional development opportunities, and an environment that fosters growth and innovation. If you are a passionate, results-driven SRE, we invite you to join us and play a pivotal role in shaping the future of our infrastructure.

Python
BS

Software Principal Engineer - Sre

Boomi Software

7+ Years Bengaluru, Karnataka, India Full-time

Position: Senior Site Reliability Engineer Join us as a Senior Site Reliability Engineer on our Reliability Team and do the best work of your career while making a profound social impact. In this role, you will design and build sophisticated systems and software that align with our customers business goals and environments. You will collaborate with product management, engineering teams, customer success, and support to deliver innovative features and enhancements across Boomi s product offerings. Key Responsibilities Incident Management & SLAs: Participate in detecting, remediating, and reporting production incidents, ensuring that SLAs and SLOs are well-defined and consistently met. On-Call Rotation: Provide on-call support for planned and unplanned events. Collaboration: Partner with engineering teams to implement improvements, standardize processes, and drive consistent results. Disaster Recovery: Lead DR exercises, game days, and readiness training with SRE and engineering counterparts. Observability & Tooling: Collaborate with service engineering teams to build and automate tooling, implement best practices in observability, and ensure the scalability and reliability of Boomi s production services. Infrastructure Automation: Automate provisioning and maintenance of Boomi s infrastructure using tools like Terraform and Ansible. Technical Mentorship: Guide and mentor other engineers through design collaboration and code reviews. What You ll Bring Essential Requirements Expertise in defining, measuring, and improving reliability metrics (SLOs, SLIs, error budgets). Strong experience in observability practices (monitoring, logging, distributed tracing), preferably using Splunk and New Relic, including the ability to create custom dashboards from scratch. Proficiency in infrastructure automation using Terraform, CloudFormation, and Ansible playbooks, with scripting experience in Python. Hands-on experience conducting and automating disaster recovery (DR) exercises in AWS, validating RPOs and RTOs. Deep understanding of AWS components and the ability to design and implement APIs for internal use. Desirable Requirements 7+ years of experience in the software engineering industry, with exposure to large-scale production systems. Cloud certification (AWS, Azure, GCP, Oracle), with experience in services such as compute, containers, and databases. Experience in containerization best practices, cloud-native concepts, and security awareness in the cloud. Working at Boomi means doing what you love, surrounded by trailblazers with an entrepreneurial spirit. Our culture fosters innovation, encourages collaboration, and celebrates the unique contributions of every individual. Take the first step toward your dream career at Boomi where ideas shape the future of technology.

AWS kubernetes Docker CI/CD terraform
II

Site Reliability Engineering Professional - Windows

Ibm India

3+ Years Bengaluru, Karnataka, India Full-time

Introduction System Engineers at IBM are integral to the company's strategic initiatives, ensuring the design, coding, testing, and delivery of cutting-edge solutions that power critical global systems. From ensuring transportation runs seamlessly to enabling secure financial transactions, the role of System Engineers is pivotal. At IBM, you ll leverage advanced development tools and work with industry-leading experts to create solutions you can take pride in. Your Role and Responsibilities We are seeking a skilled Windows Administrator to manage and maintain Windows operating systems and server networks within a critical cloud environment. In this role, you will: Install or upgrade Windows-based systems and servers. Manage user access to servers and maintain network stability and security. Troubleshoot and resolve complex IT issues. Enhance operational efficiency through automation and collaboration. This position demands a high level of technical expertise and a proactive approach to solving challenges. A successful Windows Administrator ensures seamless operations while upholding the highest security standards. Responsibilities Include: Managing critical IaaS infrastructure supporting customer workloads. Allocating and managing tools, frameworks, and assets to enhance engineering productivity and service delivery. Promoting consistent and efficient practices across teams. Who You Are You are a curious and passionate technologist eager to innovate and adopt emerging technologies. Your technical foundation is solid, and you thrive in dynamic environments, juggling multiple responsibilities to deliver integrated solutions. Who You'll Work With You will collaborate with a diverse and dynamic team, including architects, QA specialists, product managers, and delivery teams. This role promises variety, innovation, and the opportunity to make meaningful contributions daily. Required Education Bachelor s Degree Required Technical and Professional Expertise 3+ years of experience in Windows OS administration, including installation, upgrades, and troubleshooting. Expertise in setting up and configuring Windows Active Directory. 3+ years of hands-on experience with PowerShell scripting. Preferred Technical and Professional Expertise Proficiency in network configuration and troubleshooting. Experience with backup and recovery solutions. By joining IBM, you ll contribute to a collaborative environment where innovation meets practical application, driving solutions that make the world run better. Qualification : Bachelor's Degree

System Monitoring
SA

Lead/senior Devops Engineer (scripting/public Cloud)

Salesforce

7+ Years Hyderabad, Telangana, India Full-time

Lead / Senior DevOps Engineer SRE, Public Cloud & Automation (Hyderabad, India) Full-Time | Software Engineering | Salesforce Hyderabad Drive Automation. Build Scalable Systems. Shape the Future of Cloud Engineering. Join Salesforce as a Lead or Senior DevOps Engineer and play a critical role in operating one of the world s most advanced multi-cloud, microservices, and Kubernetes-based platforms. This is your opportunity to work on high-impact distributed systems that serve tens of millions of users across industries every single day. If you're passionate about DevOps, Site Reliability Engineering (SRE), infrastructure automation, and public cloud platforms, and thrive in fast-paced environments, this role is for you. What You ll Do Maintain and scale a large fleet of Kubernetes clusters powering core Salesforce services and applications Automate everything: from provisioning to deployment and monitoring using tools like Terraform, Python, Go, Spinnaker, and Puppet Ensure high availability, reliability, and performance of large-scale distributed systems Troubleshoot complex, multi-layered production issues across compute, network, and storage layers Implement self-healing systems, smart alerts, and observability solutions using tools like Grafana, Prometheus, Nagios, Zabbix Collaborate closely with engineering, architecture, and platform teams across Salesforce Evaluate and integrate emerging technologies to drive platform resilience, scalability, and automation Required Skills & Qualifications 7+ years of hands-on DevOps or Site Reliability Engineering (SRE) experience in production environments Strong background in Linux systems administration and internals Fluent in scripting or programming languages like Python, Golang, or Bash Deep knowledge of Kubernetes, Docker, and service mesh technologies Experience with AWS (Amazon Web Services) and Infrastructure as Code (IaC) tools like Terraform Strong understanding of networking concepts: TCP/IP, load balancing, switches, DNS, firewalls Familiarity with CI/CD pipelines using tools like Jenkins, Spinnaker, or similar Proficiency in using configuration management tools: Puppet, Chef, or Ansible Proven troubleshooting skills in large-scale distributed systems and cloud-native architectures A continuous learner with strong analytical thinking and collaborative communication skills Preferred / Bonus Skills Prior experience in SRE ownership of large-scale SaaS platforms Experience working with multi-region deployments and resilient architectures Hands-on knowledge of API integration with public cloud vendors (AWS, Azure, GCP) Exposure to data durability, high-availability storage systems, and clustering solutions Perks & Benefits Industry-leading benefits including healthcare, fertility support, adoption & parental leave Well-being reimbursement and generous time-off policies Continuous learning through Trailhead and world-class enablement Mentorship and leadership coaching with executive exposure Opportunities to give back via Salesforce s 1:1:1 philanthropy model Join us in redefining the DevOps landscape and making scalable, secure, and self-healing systems the foundation of modern enterprise software.

Devops scripting grafana
SS

Cloud Cost Optimizer (azure & Gcp Specialist)

Serosoft Solutions Pvt Ltd

5+ Years Indore, Madhya Pradesh, India Full-time

Cloud Cost Optimizer (Azure & GCP Specialist) Job Category: Technical Department: Infrastructure Job Location: Indore, India Experience Required: 5 Years About the Role: We are looking for a highly skilled and motivated Cloud Cost Optimization Specialist with proven experience in Microsoft Azure and Google Cloud Platform (GCP). You will be responsible for designing and implementing cost-efficient, scalable, and secure cloud strategies across a multi-cloud infrastructure, while aligning with business and performance goals. Key Responsibilities: Evaluate and optimize Azure and GCP environments for cost, performance, scalability, and reliability. Identify underutilized or over-provisioned cloud resources and suggest improvements. Develop and enforce cost governance policies and frameworks. Automate cost optimization processes using tools like Terraform, Python, and native cloud services. Collaborate with DevOps, Engineering, and Infrastructure teams to align optimization efforts with operational needs. Create and maintain dashboards and reports on cloud cost trends and KPIs. Provide architectural guidance to ensure cost-efficient workload deployment. Continuously research new Azure and GCP services, billing features, and optimization techniques. Education & Experience: 5+ years of experience in cloud infrastructure and cost optimization. Hands-on expertise with both Microsoft Azure and Google Cloud Platform (GCP). In-depth understanding of cloud services, billing models, and pricing structures. Experience with Azure Cost Management, GCP Pricing Calculator, and third-party tools like CloudHealth, Spot.io, or Apptio Cloudability. Proficient in Infrastructure as Code (IaC) tools such as Terraform or Azure ARM templates. Strong data analysis skills to interpret cloud usage and drive actionable insights. Skills & Competencies: Relevant certifications (e.g., Microsoft AZ-305, GCP Professional Cloud Architect or Cloud Engineer). Understanding of FinOps principles and cloud cost governance frameworks. Background in DevOps, Systems Engineering, or Site Reliability Engineering (SRE) is a plus. Excellent verbal and written communication skills for cross-functional collaboration. What We Offer: Learning & Growth: Support for career development and continuous learning. Innovative Projects: Be part of cutting-edge cloud and DevOps initiatives. Global Exposure: Opportunities to work on international projects. Engaging Culture: Participate in team outings, events, and celebrations. Competitive Compensation: Rewarding salary and performance-based benefits. Healthy Work-Life Balance: 5-day work week and a wellness-focused environment. Group Health Insurance: Comprehensive medical coverage for peace of mind. Open Door Policy: Your ideas and feedback are always welcome. Work from Indore: Join us in India s cleanest city with a modern, collaborative office space. Apply now and bring your expertise to a team that values innovation, efficiency, and excellence in the cloud!

Python SQL TensorFlow Pytorch
SA

Devops Engineer

Sarvam

Fresher Bengaluru, Karnataka, India Full-time

DevOps Engineer Location: Bengaluru, Karnataka, India (On-Site) Department: Engineering Employment Type: Full-Time About Sarvam.ai Sarvam.ai is a cutting-edge generative AI startup headquartered in Bengaluru, India, with a mission to make generative AI accessible and impactful for Bharat. Founded by AI experts, we are dedicated to developing high-performance, cost-effective AI agents tailored for the Indian market. We enable enterprises to tap into new opportunities, build deeper customer connections, and reshape the future of AI for India and beyond. Role Overview We are looking for a DevOps Engineer to join our team and help build and manage scalable, secure, and high-performance infrastructure. In this role, you will be a key contributor to automating deployments, managing cloud infrastructure, optimizing CI/CD workflows, and ensuring system reliability. You will work with cutting-edge technologies, including cloud platforms, containerization, and infrastructure as code (IaC), to deliver impactful solutions for AI-driven products. Key Responsibilities CI/CD Pipelines: Design, implement, and manage CI/CD pipelines for seamless software deployment and integration. Cloud Infrastructure: Deploy and manage cloud infrastructure using Terraform, Kubernetes, and Docker for scalability and high performance. Automation & Scaling: Automate infrastructure provisioning, scaling, and security compliance to support high-availability environments. Monitoring & Optimization: Implement logging, monitoring, and alerting solutions using tools like Prometheus, Grafana, ELK Stack, or CloudWatch to monitor system performance and optimize resource utilization. Security & Compliance: Enhance security and compliance by managing IAM policies, encryption, and vulnerability scanning. Troubleshooting & Root Cause Analysis: Troubleshoot system failures, perform root cause analysis, and implement improvements to ensure reliability and uptime. Collaboration: Work closely with development teams to ensure smooth deployment and operation of AI models and applications. Must-Have Skills & Qualifications Educational Background: Bachelor s degree in Computer Science, Engineering, or related field (2024/2025 graduates). Cloud Expertise: Strong experience with AWS, Azure, or GCP for deploying and managing cloud-based applications. Containerization: Proficiency in Docker and Kubernetes for building and managing containerized applications. Infrastructure as Code (IaC): Experience with Terraform, Ansible, or CloudFormation to automate infrastructure management. CI/CD Pipelines: Experience in setting up automated workflows using tools like GitHub Actions, Jenkins, or GitLab CI/CD for smooth deployments. Monitoring & Logging: Experience with Prometheus, Grafana, ELK, or similar tools to implement effective monitoring and logging solutions. Networking & Security: Strong understanding of firewalls, VPNs, SSL, and cloud security best practices for secure infrastructure. Version Control: Proficiency with Git for managing code repositories and version control workflows. Problem Solving: Strong debugging, troubleshooting, and analytical skills to resolve complex system issues. Good to Have (Preferred Experience) Serverless Computing: Exposure to serverless computing models such as AWS Lambda or Azure Functions. Message Queues: Experience with message queues like Kafka, RabbitMQ, or SQS. Site Reliability Engineering (SRE): Familiarity with SRE practices to ensure the reliability and availability of large-scale systems. Open Source Contributions: Contributions to open-source projects or a strong GitHub portfolio showcasing DevOps expertise and best practices. Impactful Work: Work on AI-driven products that are reshaping the future of technology in India. Innovative Team: Collaborate with a team of AI experts and engineers pushing the boundaries of technology. Career Growth: Opportunity to grow in a fast-growing startup at the forefront of the generative AI revolution. Cutting-edge Technologies: Work with cloud technologies, automation, and AI infrastructure to create high-impact products. Qualification : Bachelors degree in Computer Science, Engineering, or related field

CI/CD jenkins Docker kubernetes AWS
CS

Principal Cloud Development Engineer

Cloud Software Group

14+ Years Bengaluru, Karnataka, India Full-time

Job Title: Principal Cloud Development Engineer Location: Bengaluru, India About Cloud Software Group: Cloud Software Group (CSG), home to Citrix and TIBCO, is one of the largest global providers of cloud-based technologies, empowering over 100 million users worldwide. As a Principal Cloud Development Engineer, you will play a pivotal role in shaping the future of Desktop-as-a-Service (DaaS) solutions helping deliver secure, scalable, and intelligent platforms that drive modern work experiences from anywhere. We re entering an era of accelerated innovation and transformation now is the perfect time to bring your technical leadership, cloud expertise, and mentorship mindset to the forefront. About This Team: The DaaS team at CSG is responsible for designing and building scalable and resilient cloud-native microservices that power Citrix s core virtualization offerings. This team collaborates across product, architecture, operations, and customer success groups to build next-gen capabilities on Azure, AWS, and other hybrid environments. Your Role and Responsibilities: As a Principal Cloud Development Engineer, you will be expected to: Lead design and architecture discussions for cloud-native solutions within the Citrix DaaS product line. Drive the development of scalable and secure backend features, with emphasis on business logic, cloud security, and performance. Mentor junior and senior engineers, guiding them in coding best practices, design decisions, and technical growth. Collaborate with Product Managers, UX Designers, Support, and Site Reliability Engineers to build customer-centric features and maintain high service uptime. Contribute to strategic technical initiatives, including the adoption of Gen AI tools, DevSecOps automation, and performance tuning of production systems. Participate in on-call escalation support, helping debug complex issues and lead incident resolution. Promote a culture of continuous learning and improvement through code reviews, technical sessions, and post-incident analysis. Required Experience and Skills: 14+ years of experience in cloud software development using .NET (C#), Java, or equivalent Object-Oriented Programming languages. Strong computer science fundamentals (algorithms, data structures, systems design). Proven track record in building and leading cloud-native microservices with modern deployment practices (CI/CD, IaC, Kubernetes, Docker). Strong cloud platform expertise, especially in Microsoft Azure or Amazon EC2. Deep understanding of cloud security, including identity/access management, encryption, compliance, and incident response. Advanced knowledge in automation scripting (Python, PowerShell). Familiarity with troubleshooting tools like Sumo Logic, Splunk, or equivalent observability platforms. Experience with Terraform, CI/CD pipelines, and managing Kubernetes-based deployments. Strong communication, collaboration, and mentoring abilities. Preferred Qualifications: Prior experience building secure services in the DaaS, VDI, or enterprise SaaS domain. Hands-on experience with Azure Active Directory, Microsoft AD, or other identity solutions. Moderate understanding of cryptographic protocols and encryption standards. Familiarity with Agile/SAFe development methodologies. Contributions to open-source or technical publications are a plus. Impact: Influence the architecture and direction of mission-critical cloud platforms used globally. Mentorship: Be a technical leader shaping the next generation of engineers. Innovation: Work with a company at the edge of a "Cambrian leap" in cloud evolution. Culture: Inclusive, forward-thinking, and driven by curiosity and collaboration. Flexibility & Benefits: Competitive salary, performance bonus, flexible work model, health insurance, wellness programs, and more. Equal Opportunity Statement: Cloud Software Group is committed to Equal Employment Opportunity and prohibits unlawful discrimination of any kind. All qualified applicants will receive consideration without regard to race, color, religion, gender, gender identity or expression, national origin, age, disability, veteran status, or any other characteristic protected by law.

AWS kubernetes Docker terraform Python
BS

Senior Performance Engineer

Boomi Software

Fresher Bengaluru, Karnataka, India Full-time

Senior Performance Engineer Are you ready to work on world changing technologies? Today, organizations need to move with increased agility and insight to grow and thrive. Boomi is one of the hottest tech companies in the SaaS/Cloud industry, named a Leader for the eighth year in a row in the Gartner Enterprise iPaaS Magic Quadrant and recently recognized by Inc. Magazine as one of the best workplaces. Our award-winning, patented technology is transforming the world of integration by making enterprise-class integration technology accessible and affordable to companies of all sizes. Boomi provides the foundation on which your business can evolve and innovate. According to a recent survey by Vanson Bourne, connected businesses are far outpacing their competitors. We help organizations connect everything and engage everywhere across any channel, device or platform. More than 7,000 organizations are using Boomi to run better, faster and smarter. Working at Boomi means doing what you love. We hire trailblazers with an entrepreneurial spirit who can solve challenging problems, make a real impact in technology and want to build something big. If you are passionate about solving hard problems, enjoy working with world-class people and developing cutting edge technology, you should explore a career with Boomi. Learn more at http://www.boomi.com/ or visit Boomi Careers. Join us as a Performance Engineer on our Performance, Scalability and Resiliency(PSR) Engineering team in Bangalore/Hyderabad, India to do the best work of your career and make a profound social impact. What you ll achieve As a Performance Engineer, you will be responsible for validating and recommending performance optimizations in Boomi s computing infrastructure and software. You will work with our Product Development and Site Reliability Engineering teams on Performance monitoring, tuning and tooling. You will: Analyze Software Architecture (monolith and micro-service) and identify potential areas of performance, scalability and resiliency improvements Identify KPIs, perform trending and analysis, identify patterns and engineer remedial solutions for a high performant, fault tolerant and resilient platform and application stack. Design, automate and perform scalability and resiliency tests using various tools like JMeter, Chaos Monkey or similar Use observability stack to improve diagnosability and trending around Performance bottlenecks Identify performance tuning opportunities and recommend remedial solutions Take the first step towards your dream career Every Boomer brings something unique to the table. Here s what we are looking for with this role: Essential Requirements Expert in performance engineering fundamentals - arrival rate, workload models, responsiveness, computing resource utilization, time complexity, scalability, resiliency etc.. Expert in monitoring the performance using native Linux OS, Application Performance Management(APM) and Infrastructure monitoring tools Experience in analyzing crash dump, thread dump, SQL slow query log and identify performance bottlenecks Expert in recommending optimal resource configurations in Cloud, Virtual Machine, Container and Container Orchestration technologies Flexibility to work in a remote and geographically distributed team environment Desirable Requirements Experience in writing data extraction and custom monitoring tools using any programming language - Java, Python, R , Bash or similar Experience in capacity planning and modelling using AI/ML, queueing models or similar approaches Performance tuning experience in Java or similar application code

Python C++ SQL Postgresql MongoDB
HS

Java Support Engineer/software Engineer

Hsbc

Fresher Pune, Maharashtra, India Full-time

About HSBC HSBC is one of the largest banking and financial services organizations in the world, operating in 64 countries and territories. Our mission is to enable businesses to thrive, economies to prosper, and help people fulfill their hopes and ambitions. Whether you're aiming for the top of your career or exploring new directions, HSBC offers the support, opportunities, and rewards to help you realize your potential. Role: Software Engineer HSBC is seeking a Software Engineer to join our team. This role focuses on providing critical support to our production services and ensuring smooth transitions from development to production. You will assist in the daily production monitoring of Internet Banking services, collaborate across multiple teams, and help drive automation and performance improvements. If you re passionate about working in an agile and customer-first environment, this is a great opportunity to further your career. Key Responsibilities: 24x7 Support: Provide on-call support for production services, ensuring timely resolution of any issues. Service Transition: Coordinate implementation activities for a smooth transition from development to production. Production Monitoring: Handle daily production Internet Banking level 1 case checking, investigation, and provide solutions or workarounds. Collaboration: Work closely with server, infrastructure, and business teams to address technical and operational challenges. SRE & DevOps: Understand and assist in Site Reliability Engineering (SRE) and DevOps production support activities. Requirements: Technical Expertise: Strong understanding of web-based application support in a multi-tier architecture, especially within the banking domain. Familiarity with Mobile SRE support activities. Expertise in troubleshooting Java/J2EE/Microservices applications deployed on industry-standard front-end application servers such as IBM WAS, IBM WPS, JBoss, and Tomcat. Experience with UNIX systems and hands-on troubleshooting. Python-based automation experience is a plus. Proficiency in using application performance monitoring tools like AppDynamics, Splunk, BMC Patrol, HP BSM, Dynatrace, etc. Tech Stack: Knowledge of Java / J2EE, AngularJS, node.js, React, Spring, DOJO, HTML, JavaScript. Familiarity with ticketing tools such as BMC Remedy, GSD, or ServiceNow. Cloud & Systems Knowledge: Familiarity with cloud management, Oracle databases, LDAP, MQ, TCP/IP networks, web servers, and data center management. Understanding of Microservices, REST APIs, and integration with platforms like Mule Gateway, AnyPoint, and PCF. Experience in application migration and a clear understanding of the process. Preferred Qualifications: Experience with business-critical server application support. Background working in cloud environments. Understanding of application migration processes and related activities. At HSBC, you ll join a global team working at the forefront of technology and banking services. We offer opportunities for growth, support for work-life balance, and the ability to make a meaningful impact on the global economy. If you re ready to take on a dynamic, customer-facing role and thrive in a fast-paced, collaborative environment, we want to hear from you! Qualification : Candidates with good understanding of Cloud, Oracle database, LDAP, MQ, TCP/IP Networks, Webservers, Data Centre will be preferred.

SQL CI/CD Docker kubernetes
BS

Devops Engineer

Bmc Software

3+ Years Pune, Maharashtra, India Full-time

Company Overview: BMC Software is an award-winning, equal opportunity employer that fosters a diverse and culturally rich work environment. The company is committed to giving back to the community, and innovation is at the heart of everything we do. We create an atmosphere where your contributions are celebrated, and your ideas are heard. Our SaaS Ops department focuses on delivering exceptional SaaS experiences to our customers by utilizing cutting-edge technologies. We continuously strive to grow by adopting the latest innovations, and we offer a global, versatile environment where professionals can thrive. Role Overview: As a DevOps Engineer, you will join our dynamic SaaS Ops team to design, develop, and implement complex enterprise applications using the latest technologies. You will play a key role in driving the adoption of DevOps processes and tools across the organization. This position offers exciting opportunities to work with industry-leading tools and practices, contributing to the growth and success of both BMC and your own professional development. Responsibilities: End-to-End Product Development: Participate in all aspects of SaaS product development, from requirements analysis to product release and ongoing sustenance. Ensure the delivery of high-quality enterprise SaaS solutions within the specified schedule. DevOps Process Adoption: Drive the adoption of DevOps processes and tools throughout the organization. Develop and maintain Continuous Delivery Pipelines to optimize the deployment process. Technology Integration: Learn and implement cutting-edge technologies to build enterprise SaaS solutions at scale. Work with cloud technologies, containerized environments, and automation tools to enhance application performance and reliability. Collaboration: Collaborate with cross-functional teams, including R&D, Operations, Support, and others, to ensure seamless integration of DevOps practices into the workflow. Documentation & Troubleshooting: Design and document Standard Operating Procedures (SOPs), architecture artifacts, and design documents. Use troubleshooting skills to address issues across different platforms, ensuring minimal downtime and optimal performance. Required Skills & Qualifications: Experience: 3+ years in a software engineering function, preferably with experience in DevOps, SaaS, and automation. Technical Expertise: Strong experience with CI/CD pipelines, containerized deployments, and maintaining production environments. Proficiency in automation scripting languages such as Python, Groovy, Ansible, or Shell scripting. Hands-on experience with Jenkins, Docker, Helm, Git, Terraform, and Jira. Knowledge of Web service protocols (REST, JSON) and experience working with Relational Databases (e.g., PostgreSQL, MS SQL). Containerization & Cloud Technologies: Familiarity with Kubernetes (PODs, persistent storage, ingress, routes) and cloud deployment models (public, private, hybrid). Exposure to ElasticSearch, Grafana, Prometheus, and other monitoring tools. Site Reliability Engineering (SRE): Understanding of SRE principles and their implementation for SaaS services to ensure scalability, performance, and reliability. Operating Systems: Proficient working on Windows and Linux platforms. Agile Methodology: Experience working in an Agile environment with cross-functional teams. Soft Skills: Excellent troubleshooting, communication, and collaboration skills. Hardworking, dedicated, and capable of handling time-sensitive deadlines. Education: Bachelor s degree in IT or a related field, or equivalent professional experience. Bonus Skills (Nice-to-Have): Familiarity with BMC Helix products (ITSM, Digital Workplace, Helix Platform) is an advantage. Previous experience in Site Reliability Engineering (SRE) for SaaS products will be a plus. Work Schedule & Benefits: This position may require occasional weekend work during scheduled production activities and after-hours work as needed. As part of BMC's commitment to equal opportunity, employees benefit from a supportive and inclusive culture, with opportunities for professional growth. Compensation & Rewards: The midpoint of the salary band for this role is 1,638,100 INR. Actual salary will depend on factors such as skills, experience, certifications, and other business needs. BMC offers a comprehensive compensation package, including a variable pay plan and country-specific benefits. Why Join BMC? BMC is a company that thrives on innovation, collaboration, and a commitment to creating a work environment that allows you to bring your best self to work every day. If you re looking for a place where you can make an impact, work with cutting-edge technology, and grow alongside talented professionals, this is the place for you. Be yourself at BMC, and help us shape the future of SaaS! Qualification : Bachelors degree in IT or a related field, or equivalent professional experience.

ST

Power Electronics Engineer

Solaredge Technologies

4+ Years Bengaluru, Karnataka, India Full-time

Job Description Power the Future with us! SolarEdge (NASDAQ: SEDG), is a global leader in high-performance smart energy technology, with over 5000 employees, offices in 34 countries, and millions of products installed in over 133 countries. Our diverse product offering comprises intelligent solar inverters, battery storage, backup systems, EV charging, and complete home energy management ecosystems. By leveraging world-class engineering capabilities and with a relentless focus on innovation, we strive to create a world where clean, green energy from the sun is the primary source of power for our homes, businesses, and just about everywhere we thrive. Our R&D division is growing globally, and we are looking for an experienced Power Electronics Engineer to join our dynamic team at the new R&D site in Bangalore, India. As a Power Electronics Engineer at SolarEdge India R&D, you will play a pivotal role in the design, development, and optimization of power electronics and power electronics systems for our advanced solar energy products. You will be responsible for driving the innovation and technical excellence of our power solutions, contributing to the success of SolarEdge's mission to make solar energy more accessible and efficient. Responsibilities: Design, analysis, and development of advanced power electronics and power systems for SolarEdge's solar energy products, including inverters, power optimizers, and energy storage solutions. Collaborate with cross-functional teams, including electrical engineers, Mechanical Engineers/Designers, PCB Layout Engineers, and firmware developers, to ensure seamless development, integration, and optimization of power systems. Conduct power system studies, such as load flow analysis, transient stability, and harmonic analysis, to assess system performance and reliability. Perform detailed design and analysis of power electronic circuits, ensuring compliance with industry standards and safety regulations. Prepare detailed design documentation, Schematics, BoM, and test procedures. Lead the testing and verification of power electronics and power systems, both in the lab and field, to ensure they meet design specifications and quality standards. Participate in design reviews, providing technical expertise and guidance to the team to drive continuous improvement and innovation. Collaborate with suppliers and manufacturing teams to support the transition of designs from R&D to mass production, addressing any design-related issues during production. Mentor and guide junior engineers, fostering a collaborative and innovative work environment. Job Requirements Bachelor s (B.E/B.Tech) or master s (M.E./M.Tech) degree in electrical /electronics Engineering with a specialization in Power Electronics or Power Systems. 4+ years of hands-on experience in power electronics design and power system analysis, preferably in the solar energy or renewable energy industry. Strong understanding of power semiconductor devices (Including SiC and GaN), gate drive circuits, and magnetic components used in power converters. Experience with simulation tools (e.g., PSpice, LT SPICE, Simulink etc.) Knowledge of power converter topologies (e.g., DC/DC, DC/AC and AC/DC), including resonant and bi-directional converters and grid-tied inverters. Excellent knowledge of PCB layout rules, considering high-current traces, thermal management, creepage, clearance requirements for HV and LV traces. Ensure compliance with best practices for power distribution, component placement, and impedance control. Implement EMI/EMC best design practices to ensure compliance with regulatory standards. Familiarity with international safety and regulatory standards for power electronics and solar energy products. Excellent problem-solving skills and the ability to troubleshoot and resolve complex technical issues. Strong communication and interpersonal skills to work effectively in a cross-functional team environment. Proven track record of delivering high-quality power electronics designs from concept to production. Results-oriented mindset with a focus on achieving tangible and measurable results. Qualification : Bachelors (B.E/B.Tech) or masters (M.E./M.Tech) degree in electrical /electronics Engineering with a specialization in Power Electronics or Power Systems.

No results found

Modify search criteria or create an alert to get relevant jobs as soon as they’re posted

Create an alert

1 - 20 of 3182

Continue to Save

Please login to your jobseeker account, or create a new one to save this job.

Filter jobs

Feedback

Share Feedback