Overview
Salary: $70.32-78.13 Hourly
Aquent is partnering with a leading financial services institution that is revolutionizing how individuals manage their financial futures. This organization is at the forefront of integrating cutting-edge technology, including AI/ML, to enhance the reliability and performance of its critical applications. By joining this team, you will play a pivotal role in shaping the future of financial technology, driving innovation, and ensuring seamless, high-availability experiences for millions of users. Your contributions will directly impact the stability, efficiency, and intelligence of the platforms that power financial success. Unleash Your Expertise as a Site Reliability Engineer Are you a skilled engineer passionate about combining software systems engineering with robust operations, especially through AI/ML-driven approaches? We are seeking an innovative Site Reliability Engineer to join a dynamic team dedicated to managing and optimizing large enterprise and mission-critical applications. This is an exciting opportunity to evangelize the SRE mindset, build groundbreaking tools, and implement intelligent automation that significantly reduces manual toil and elevates operational throughput. You will be at the heart of designing and deploying advanced AI/ML-driven automation pipelines, enhancing observability, and creating proactive operational response systems that set new industry standards for reliability. What You'll Do
- Champion the SRE mindset and drive problem-solving through systematic approaches and innovative solutions.
- Identify and seize opportunities to develop unique tools and resolve complex operational challenges within critical enterprise applications.
- Architect and implement production automation solutions that measurably reduce manual effort and boost operational efficiency.
- Design and deploy AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems, including anomaly detection and predictive alerting, to elevate platform reliability.
- Lead the expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
- Collaborate extensively with Engineering, Scrum, and Operations teams, providing crucial technical expertise and support for key initiatives focused on system availability and reliability.
- Efficiently triage alerts, diagnose, and resolve critical issues, managing change implementations with clear communication and minimal risk.
- Develop essential tools, frameworks, and instrumentation to validate and enhance the success of application rollouts, leveraging AI/ML capabilities for operational visibility and validation at scale.
- Advocate for AIOps platform adoption and ML-assisted observability practices across the team.
- Coordinate robust capacity planning through data-driven trend analysis and ML-informed forecasting.
- Develop CI/CD orchestration systems to streamline software delivery to production, championing GitOps concepts and AI-assisted pipeline optimization.
- Perform real-time troubleshooting of mission-critical application workflows, integrating feedback directly into product development cycles.
- Participate actively in on-call support rotations, ensuring continuous operational excellence.
What You'll Bring
- 6-8 years of hands-on experience in enterprise-level administration and support.
- 6-8 years of experience crafting automation scripts, developing application dashboards for proactive monitoring, and configuring alerts for early issue detection.
- 6-8 years of practical experience with SDLC and process improvement methodologies.
- Extensive hands-on experience with enterprise systems administration, monitoring, and deployment activities.
- Proficiency with Windows 2019/2022 and Linux operating systems hosted via Virtual Machine.
- Experience in Cloud application configuration, deployment, support, and migration.
- Solid understanding of IP networking fundamentals, including DNS, DHCP, firewalls, and IP routing.
- Familiarity with large-scale distributed systems and high-availability architectures.
- Expertise in Linux and Windows system administration, troubleshooting, and performance tuning.
- Development experience in one or more programming languages such as .NET, PowerShell, Java, Python, or Bash.
- Knowledge of one or more database systems, including SQL, Oracle, or MongoDB.
- Working knowledge of Actimize.
- Familiarity with one or more Message Brokers, such as Solace, RabbitMQ, IBM MQ, or Kafka.
- Experience with Splunk, AppDynamics, or similar observability tools.
- Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
- A Bachelor's degree in computer science or a related discipline.
- Strong customer orientation with an affinity for proactively owning, communicating, and following through on projects and issues.
- An extreme sense of ownership to meticulously resolve problems within a distributed environment.
- A gritty resolve to delve deep into technical issues within a complex login ecosystem.
- A self-starter mentality with the ability and confidence to independently resolve issues and deliver results to the team.
Bonus Points
- Experience within the financial services industry.
- Proficiency with Agile methodologies.
- Hands-on experience with AIOps platforms or ML-driven observability tooling.
- Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
- Familiarity with CI/CD tools (e.g., Harness, Jenkins, GitHub Actions) or GitOps concepts.
- Exposure to container orchestration platforms (e.g., Kubernetes, OpenShift) or cloud platforms (e.g., AWS, Azure, GCP).
About Aquent Talent: Aquent Talent connects the best talent in marketing, creative, and design with the world's biggest brands. Our eligible talent get access to amazing benefits like subsidized health, vision, and dental plans, paid sick leave, and retirement plans with a match. Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We're about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive. #LI-LP1
|