高级基础设施工程师:可观测性
Senior Infra Engineer: Observability
职位描述
Railway的核心使命是让软件工程师更具效率。我们认为,人们应该获得强大的工具,这样他们就可以花更少的时间进行设置,更多的时间进行实际工作。
许多基础设施平台只关注如何部署你的单个应用程序,而没有关注这些应用程序如何协同工作。诸如“如何构建零停机时间的部署系统”、“如何实现服务间通信”等问题通常由工程师自行定义。
在Railway,我们的目标是为所有这些问题提供一个全面的解决方案。因此,在定义我们的网络基础设施时,我们特别注重细节。
注意:网络属于平台工程范畴。如果你有相关专长,我们很乐意与你交流!但同时也要说明,你可能会做很多非网络和平台相关的工作。
“如果更多的工程师像我一样讨厌技术,世界会变得更美好。我设计的东西,如果成功的话,没人会注意到。一切都会正常运行,并且是自我管理的”
- Radia Perlman
关于该职位
在此职位中,你将:
1. 构建摄入管道,以处理每秒百万次以上的日志、指标和其他遥测数据流
2. 构建可扩展、容错的实时警报引擎,用于通知用户阈值违规情况
3. 设计丰富的后端可观测性API,与产品团队合作,为用户即时理解其应用提供出色体验
4. 提供API以访问实时日志/指标流,供仪表板和产品团队使用
5. 从零开始构建Golang/Rust GRPC服务,能够支持数万名用户,以及即将到来的百万级用户
6. 定义可以使用不可变基础设施原则通过Terraform和Ansible进行销毁、故障转移和重新构建的基础架构
7. 编写工程需求文档,将想法转化为定义的任务,再到实现,最后监控其成功
8. 与我们的TypeScript和GraphQL边缘服务交互,为内部和潜在外部使用暴露你的微服务API
这是一个影响深远、自主性强的职位,对公司的文化、发展轨迹和成果有直接影响。
关于你
- 对分布式系统有深入的理解。你喜欢构建容错、弹性和可扩展的服务
- 对VictoriaMetrics、ClickHouse和其他用于构建系统的工具感兴趣
查看英文原文
Job description
Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.
Many infrastructure platforms simply focus on how you deploy your singular application, and now how these applications function in concert. Questions like “How do you build systems for zero downtime deployment”, “How do you do service-to-service communications”, etc are usually left up to the engineers to define.
At Railway, our goal is to be an all encompassing solution to all these problems. As such, we take special care as we define our networking infrastructure.
Note: Networking falls under the platform engineering umbrella. If you’re specialized, we’d love to chat! That said, we’d also like it noted you’re probably going to do a lot of non-networking + platform things
“But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing”
- Radia Perlman
ABOUT THE ROLE
For this role, you will:
1. Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
2. Build scalable, fault tolerant alerting engines for notifying users, in real-time, of threshold breaches
3. Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
4. Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
5. Build Golang/Rust GRPC services from scratch capable of supporting tens of thousands of users, and the million+ to come.
6. Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible.
7. Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success.
8. Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption
This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome.
ABOUT YOU
- A strong understanding of distributed systems. You enjoy building fault tolerant, resilient, and scalable services
- Interests in VictoriaMetrics, ClickHouse, and other systems for building observability stacks from the ground up
- A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo.
- The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around
- A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
- A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
- A great set of communication skills for getting your point across, solution implemented, and beyond
We value and love to work with diverse persons from all backgrounds
Things to know
For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages.
- We're distributed ALL across the globe, and that's only going to be more and more distributed. As a result, stuff is ALWAYS happening.
- We do NOT expect you to work all the time, but you'll have to be diligent about your boundaries because the end of your day may overlap with the start of someone else's.
- We're a small team, with high ownership, who are not only passionate about what we do, but seek to be exceptional as well. At the time of writing we're 21, serving hundreds of thousands of users. There's a lot of stuff going on, and a lot of ambiguity.
- We want you to own it. We believe that ownership is a key to growth, and part of that growth is not only being able to make the choices, but owning the success, or failure, that comes with those choices.
Benefits and perks
At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page https://railway.app/careers/full-stack.
Beyond compensation, there are a few things that we believe that make working at Railway truly unique:
- Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work.
- Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company.
- Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions.
- Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there.
How we hire
No tricks. No surprises. Here's the entire process:
1. Talk with us about the role
- This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go.
2. Work on a small project to discuss in the interview
- Asynchronously implement the following:
1. Imagine a theoretical or actual system like Railway which can manage stateless and stateful compute workloads. Design the engine for managing observability
2. Interview Structure (60 Minutes):
3. Pre-work (before your interview): Complete your solution (advised)
4. 0-5m: introduction
5. 5-50m: Building (or expanding) your solution
6. 50-60m: Questions on Railway/Tech/etc.
You can, and SHOULD! ask us questions ahead of time. Ask away!
3. Review your solution with the Team
You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team.
Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution.
1. Interview Structure (60 Minutes):
1. Prework (submitted before your interview): Complete your solution
2. 0-5m: introduction
3. 5-50m: Building (or expanding) your solution
4. 50-60m: Questions on Railway/Tech/etc
4. Meet the Team
1. You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company.
1. Looking for: How you work with the rest of the team and communicate.
5. Offer and Details Chat with CEO
1. Finally, we will go over the process, the role, and hammer out the details about your position, onboarding, and all the deets.
Final Note: The interview goes both ways. Once again, please ask us things. Many things! Hard things. That's what we're here for.