站点可靠性工程师
Site Reliability Engineer
Tyk 是一家什么样的公司,我们做什么?
Tyk API 管理平台正在推动连接世界并赋能新产品和服务。我们正在改变组织连接任意数量系统和服务的方式。无论是内部、外部、公开还是高度加密的系统,Tyk 都能帮助企业在零售、金融、电信、医疗或媒体等行业(仅举几例)创造价值!
如果你在线银行、使用应用查看新闻,甚至驾驶联网汽车,API 以及由此延伸的 Tyk,让这一切成为可能。自 2015 年成立以来,我们在伦敦(英国)、伦敦(安大略省)、亚特兰大和新加坡设有办公室,我们的 B2B 平台在全球拥有数十万用户。使用 Tyk 的品牌包括乐天、贝尔、T-Mobile、RBS、Capital One 和 Vinci。我们的用户群体来自世界各地——甚至南极洲。
我们的使命
Tyk 的使命是连接世界上每一个系统。我们从构建 API 管理平台开始。
完全灵活,默认远程,强调责任
我们为所有员工提供无限制的带薪假期,并允许全球任何地方远程办公。为什么?Tyk 成立之初就秉持着为员工提供灵活性和自主权的原则,我们认为这能让员工取得最佳成果。这也意味着我们可以打造最优秀的团队,地点和工作时间不再是障碍。
如果这听起来像是一个适合你的环境,请继续阅读以了解更多信息。
职位描述:
我们正在寻找一名站点可靠性工程师,负责管理、维护、改进和支持我们的平台。你天生充满好奇心,总是寻找提升的方法,我们会期待你提出新的想法、解决方案和指标来提升平台。你还将是我们客户事件管理的第一线人员,并将帮助制定未来的应对策略。这是成为 Tyk 重要一员的绝佳机会,随着我们继续前进。
作为一家以远程优先的公司,你将有机会与行业领先的分布式团队合作。能够接触到全球各地的专业知识,将为你提供支持和机会,不仅塑造 Tyk 的云平台,也塑造整个 Tyk,因为我们持续成长。
要求
以下是你的职责:
- 在符合 SL(A/I/O)s 的前提下维护全球 Tyk Cloud
- 识别可靠性问题并与团队协作解决
查看英文原文
Who are Tyk, and what do we do?
The Tyk API Management platform is helping to drive the connected world and power new products and services. We’re changing the way that organisations connect any number of their systems and services.Whether internal, external, public or highly encrypted systems, Tyk helps businesses drive value across the retail, finance, telecoms, healthcare, or media industries (to name just a few!)
If you’ve banked online, used an app to check the news, or perhaps even driven a connected car, API’s, and by extension, Tyk, make that possible. Founded in 2015 with offices in London – UK, London – Ontario, Atlanta and Singapore, we have many thousands of users of our B2B platform across the globe. Brands using Tyk range from Lotte, Bell, T Mobile, to RBS, Capital One and Vinci. We have a varied user base hailing from every continent – even Antarctica.
Our Mission
Tyk is on a mission to connect every system in the world. We’ve started by building an API Management platform.
Total flexibility, default remote, radical responsibility
We offer unlimited paid holidays and remote working from anywhere in the world, for everyone, Why? Tyk was founded on the principle of offering flexibility and autonomy to our employees, we believe this allows our employees to achieve their best results. It also means we can build the best possible team, location and working hours are no barrier.
If this sounds like an environment that you believe could work for you then read on to find out more.
The role:
We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform. You will be curious by nature, always looking for ways to improve, as we will look to you for new ideas, solutions and metrics on how we can improve the platform. You will also be our first line of incident management to our clients and will help define our response going forward. This is a great opportunity to become an integral part of Tyk as we continue on our journey.
As a remote first company, you will have the opportunity to work with an industry leading distributed team. Having access to expertise from across the globe will give you both the support and opportunity to help shape not only Tyk’s Cloud platform but also the Tyk as a whole as we continue to grow.
Requirements
Here’s what you’ll be responsible for:
- Maintaining global Tyk Cloud within SL(A/I/O)s you will help to define
- Identifying reliability issues and working together with your squad to solve them
- Identifying and introducing new metrics and building relevant dashboards
- Participating in the on-call rotation
- Working with your squad to expand multi-region and multi-cloud reach of the platform
- Documenting operational knowledge
- Conducting post-incident analysis
- Automating common tasks
- Be a key shaper and contributor to our continuous improvement agenda – be it the clarity of our user stories, how we estimate, communicate with other teams or customers – we expect this role to be advocate of continuous improvement
- Reliability of our new global Tyk Cloud platform
- Automation of operations and support
- Writing and maintaining documentation on SRE processes and policies
- Recommending and implementing ways of driving operational efficiency and driving down our cost to run, without impacting service
- Assisting in penetration testing for Cloud through liaising with our provider, providing technical details, and environment setup
- Incident management
Here’s what we’re looking for:
Experience
- Strong collaboration skills
- Launching and operating production scale kubernetes clusters
- Designing and operating infrastructure on AWS and other providers
- Operating MongoDB (or other document database) clusters
- Operating Redis (or other key-value storage) clusters
- Administering Linux servers
- Maintaining distributed software
- Operating Prometheus and Grafana
- Operating logging collection and analysis systems
- Participating in the on-call rotation(16:00pm – 4:00am UTC)
Skills:
- Kubernetes & containers (advanced)
- AWS / EKS (advanced)
- Linux (advanced)
- Go (proficient)
- Terraform and IaC in general (proficient)
- Helm (proficient)
- MongoDB (or similar)
- Redis (or similar)
- Monitoring – prometheus, grafana, thanos (familiar)
- Grasp of networking concepts (subnets, routing, peering, load balancing, NAT, etc.)
- Common networking protocols (DNS, TCP/IP, HTTP, TLS, UDP)
- Proactive, energetic, innovative and change oriented
Nice to have:
- GCP or Azure
- Bare metal infrastructure engineering
- API management experience
- Large scale distributed storage management
- Familiarity with Rancher
- CKA/CKAD/CKS
- Creating and delivering production software in Go language
Benefits
Here’s why you should join us:
- Everyone has unlimited paid holiday.
- We have total flexibility in hours, as we believe creativity flows better when our people are given freedom to decide when they are most productive. Everyone is unique after all.
- Employee share scheme
- Generous maternity and paternity leave
- Company retreats
We all share the same vision – we value authenticity, respect, responsibility, independence, honesty, diversity and inclusion and most importantly treating others how you wish to be treated. We look for like-minded people who bring their personalities to work everyday, strive to achieve their personal goals and who are willing to challenge the way we do things, why? – to make what we do even better!
Our values tell the story of Tyk – here’s how:
· It’s ok to screw up!
We’ve found that it’s often the ‘stupid’ or unexpected ideas that turn out to be the successful ones – so try it, at least we can say we have!
· The only stupid idea, is the untested one!
It’s in our DNA – starting a business with founders 12 hours apart, giving our gateway away for free – sure, we did that, and we’d do it again!
· Trust starts with you – make it count!
Trust is a two-way street – instill it from day one!
· Assume best intent!
We have each other’s back – we’re all on the same team. Think before you speak or act.
· Make things, better!
Always try to leave things better than when you found them – change is constant, inevitable and embraced! Be that change we want to see.
What’s it like to work here?! check it out:
Tyk is an equal opportunities employer and we are determined to ensure that no applicant or employee receives less favourable treatment on the grounds of gender, age, disability, religion, belief, sexual orientation, marital status, or race, or is disadvantaged by conditions or requirements which cannot be shown to be justifiable.
You can see more about us here
Originally posted on Himalayas