系统可靠性工程师,技术负责人
Site Reliability Engineer, Tech Lead
你是否想加入一家创新的物流科技公司?
Loadsmart是一家估值超过10亿美元(真正的科技独角兽)的成长型科技公司!
我们是一群行业资深人士和以用户为中心的工程师,利用创新技术勇敢地重新定义货运的未来,帮助发货人、经纪人、仓库和承运商用更少的资源完成更多运输。
总部位于芝加哥,团队分布在全球各地,Loadsmart持续吸引致力于推动有意义变革的顶尖人才。我们寻找践行以下核心价值观的专业人士:好奇心、清晰度、成果、承诺和团队合作。
在SRE、技术主管的角色中,你将构建和维护公司的内部平台,推动运营卓越,并赋能整个工程团队。你应该有分析、提出并实施更安全系统和流程的经验。与平台工程中的工程小组紧密协作,确保我们的应用既安全又可靠。作为一位亲力亲为的领导者,你将积极参与技术工作,同时与内部利益相关者和组织内各工程小组密切合作,确保我们的应用既安全又可靠。
部门:工程
地点:巴西任何地方 - 远程
你将负责的工作
- 与我们富有创造力、关系紧密的开发团队合作并提供支持。
- 设计、部署和运行Loadsmart的关键系统,同时平衡可靠性、成本和敏捷性。
- 在工程团队中发挥关键作用,推动可靠性项目。
- 利用你的直觉解决问题的能力和积极向上的态度,解决具有挑战性和令人兴奋的问题,激励身边的人。
- 收集指标并理解其业务影响,鼓励团队也这样做。
- 对系统运行问题进行故障排查和根本原因分析。
- 对平台的服务级别协议和目标负责。
- 在需要时提供非工作时间的基础设施支持。
- 负责软件基础设施项目。
- 通过代码和规范评审,提供、接受和给予建设性的反馈。
- 熟悉AI代理和代理工作流,将AI应用于整个软件开发生命周期(AI辅助编码)、大语言模型、MCP服务器/网关,以及新兴AI工具如何提升可靠性和运营,是加分项。
所需资格:
- 1-3年领导可靠性工作的经验
查看英文原文
ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?
Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!
We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.
With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.
In the SRE, Tech Lead role you will build and maintain the company's internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.
DEPARTMENT: Engineering
LOCATION: Anywhere in Brazil - Remote
WHAT YOU GET TO DO
- Collaborate with and support our creative, tight-knit development team.
- Design, deploy, and operate Loadsmart's critical systems while balancing reliability, cost, and agility.
- Play a key role in driving reliability projects with engineering teams.
- Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
- Collect metrics and understand their business impact, encouraging the team to do the same.
- Perform troubleshooting and root-cause analysis of system operation issues.
- Be accountable for the platform's Service Level Agreements and Objectives.
- Provide infrastructure support during off-hours as needed
- Take ownership of software infrastructure projects
- Seek, give, and receive constructive feedback through code and specification reviews.
- Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.
REQUIRED QUALIFICATIONS:
- 1-3 years leading Reliability Work across multiple engineering squads
- Over 5 years of experience in Cloud Computing, SRE/DevOps
- Proven experience collaborating with internal stakeholders across multiple engineering squads
- Strong project management skills with a demonstrated ability to delegate and mentor team members
- Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
- Detail-oriented with high initiative and self-motivation
- Strong understanding of software engineering principles and how systems work under the hood
- In-depth knowledge of modern networking and operating systems
- Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
- Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
- Solid troubleshooting and system engineering experience in UNIX/Linux production environments
- Experience with monitoring, alerting, and incident management
- Proficiency in automating tasks with scripting languages like Python, Bash, etc
- Experience or exposure to PostgreSQL and DBA responsibilities is a plus
- Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.
WORKING AT LOADSMART:
• Competitive base salaries - we believe in rewarding top talent
• Extremely competitive Equity package - become a shareholder in our company!
• Loadie Time Off - PTO and sick days without a limit
At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.
It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates' specific locale.
Originally posted on Himalayas