系统可靠性工程师 (欧洲)
Site Reliability Engineer (Europe)
Site Reliability Engineer
关于 Arango:
Arango 提供了一个统一的、原生多模型上下文数据平台,为 AI 代理、助手和应用程序提供统一、实时且可信的业务上下文,以实现大规模推理、决策和行动。
Arango 上下文数据平台通过简化的架构将分散的企业数据与 LLM、copilots 和 AI 代理连接起来。通过在一个平台上结合图、向量、文档、键值和搜索功能,Arango 消除了许多组织为实现企业 AI 而构建的复杂堆栈。
NVIDIA、HPE、伦敦证券交易所、PSI CRO、美国空军、NIH、西门子、Matpriskollen 和 Articul8 等组织都信任 Arango,帮助企业在 AI 试点阶段快速转向可靠的生产系统,同时降低基础设施复杂性和总拥有成本。Arango 是 NVIDIA Inception 计划和 AWS ISV Accelerate 计划的成员。了解更多请访问、LinkedIn 和 G2。
地点:(远程)- 印度
职位概述:
在 ArangoDB,我们正在构建一个强大的云原生基础设施,以支持我们的分布式数据库系统,这些系统为多个行业的关键任务应用提供支持。我们正在寻找一名 Site Reliability Engineer (SRE),以确保我们基础设施和应用的可靠性、可扩展性和性能,重点关注自动化、监控和优化云环境。
作为 Site Reliability Engineer (SRE),您将负责维护和提升运行在 Kubernetes 和云环境(AWS、Google Cloud)上的分布式数据库系统的可靠性。您将设计、实施和维护可扩展的基础设施解决方案,改进并扩展这些解决方案的可观测性,并解决复杂的系统问题。期望您能够来工作,用 Golang 编写干净高效的代码,并与开发团队紧密合作。
您的目标是确保我们基于云的系统的高可用性和性能,自动化重复任务,并增强我们的 CI/CD 流水线。如果您热衷于构建弹性系统、管理云基础设施,并使用 Golang 创建可扩展的解决方案(或有意愿学习 Golang),我们希望听到您的声音!
主要职责:
● 在 AWS 和 Google Cloud 平台上设计、实施和维护云基础设施。
● 确保可扩展性和性能。
查看英文原文
Site Reliability Engineer
About Arango:
Arango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.
The Arango Contextual Data Platform connects fragmented enterprise data with LLMs, copilots, and AI agents through a simplified architecture delivered out of the box. By combining graph, vector, document, key-value, and search capabilities in a single platform, Arango eliminates the complex stacks many organizations build to operationalize enterprise AI.
Trusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air Force, NIH, Siemens,, Matpriskollen, and Articul8, Arango helps enterprises move from AI pilots to reliable production systems faster while lowering infrastructure complexity and total cost of ownership. Arango is a proud member of the NVIDIA Inception Program and the AWS ISV Accelerate Program. Learn more at, LinkedIn, and G2.
Location: (Remote)- India
Job Overview:
At ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments.
As a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams
Your goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!
Key Responsibilities:
● Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
● Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
● Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
● Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
● Develop strategies for disaster recovery, high availability, and fault tolerance.
● Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
● Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
● Participate in on-call rotations to support critical production systems and respond to incidents.
● Collaborate with cross-functional teams to improve overall system reliability and scalability.
● Collaborate with the Customer Success team to resolve customer issues.
Required Skills and Qualifications:
● Experience: SRE or DevOps Engineer background in cloud-native environments. Self-organized, autonomous remote team player with strong communication skills.
● Cloud & Infrastructure: AWS and GCP; advanced Linux internals (processes, environment variables); containerization and orchestration (Docker, Kubernetes at scale).
● CI/CD & Observability: CI/CD pipelines (Jenkins, CircleCI); monitoring, alerting, and logging (Prometheus, Grafana, ELK stack); Git version control.
● Networking & Security: Core networking, security best practices, and systematic troubleshooting of complex infrastructure issues.
● Development: Programming proficiency in Golang or Python.
Nice-to-Have:
● Experience managing distributed databases or large-scale data storage systems. ●Knowledge of security best practices in cloud environments.
● Experience with scripting languages like Python or Bash.
● Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. ●Experience working with GitOps
● Strong programming skills in Golang, with experience in developing automation tools, scripts, or services.
What Makes Arango Special?
At Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.
Working at Arango means:
Contributing to cutting-edge AI and data infrastructure
Collaborating with experienced engineers, marketers, and product leaders
Helping shape how enterprises build AI-powered applications
If you're excited about the intersection of AI, data, and social media, we’d love to hear from you.Originally posted on Himalayas