10倍软件工程师
10x Software Engineer
Lucidya 是中东和北非地区领先的 AI 驱动的客户体验管理(CXM)平台,使企业能够跨数字渠道大规模地了解、互动并服务客户
随着我们向上市规模迈进,我们正在重建部分平台以实现以下目标:
- 极高的可靠性(基于 SLO 的工程)
- 高规模分布式处理(数十亿的数据点)
- 原生 AI 架构(LLM + 实时智能)
我们不追求头衔,我们追求影响力!无论你是初级还是高级工程师,重要的是你比普通工程师更快地行动和更深入地思考。我们以极端的所有权文化运作。问题不属于某个团队或工单,而是属于第一个发现它的人。这个职位的合适人选不会对故障系统视而不见。如果某处出现故障,你会去修复它,即使它不是“你的工作”
#### 你将参与的工作
#### 微服务与分布式架构
- 在处理数十亿数据点的 100 多个微服务生态系统中设计和运营高吞吐量、事件驱动的流水线
- 使用 RabbitMQ 构建和扩展分布式消息系统,包括背压管理、消费者扩展和队列健康监控
- 开发和维护具有高级路由功能(多上游、流量拆分、环境隔离)的 API 网关层
- 为企业客户设计 SSO 和身份联合,支持多 IdP 路由,并与核心服务零耦合
- 定义跨越 Ruby 和 Python 的采集、处理和交付流水线中的清晰服务边界
#### 性能与系统
- 诊断和解决复杂的生产问题(例如死锁、队列耗尽、连接池饱和)并消除根本原因
- 优化 PostgreSQL 以应对大量写入工作负载,包括争用管理、模式设计、触发器和连接扩展
- 设计和调优 Elasticsearch 以实现大规模搜索、索引和实时阿拉伯语相关性
- 根据工作负载特性在多进程和异步架构之间做出明智的权衡
#### 可观测性与可靠性
- 使用 Grafana、Loki、分布式追踪和 SLOs 在大规模系统中构建和维护可观测性
- 全程负责生产事件,追踪跨队列、搜索系统和外部集成的故障
- 领导多服务流水线的根本原因分析并实施预防措施
- 构建内部工具以提高工程效率和自动化
查看英文原文
Lucidya is a leading AI-powered Customer Experience Management (CXM) platform in the MENA region, enabling enterprises to understand, engage, and serve customers across digital channels at scale.
As we move toward IPO-scale, we are rebuilding parts of our platform to achieve:
- Extreme reliability (SLO-driven engineering)
- High-scale distributed processing (billions of data points)
- AI-native architecture (LLM + real-time intelligence)
We don’t optimize for titles, we optimize for impact! Whether you are junior or senior doesn’t matter and what matters is your ability to move faster and think deeper than the average engineer. We operate with extreme ownership. Problems don’t belong to teams or tickets where they belong to whoever sees them. The right person for this role doesn’t walk past a broken system. If something is failing, you fix it even if it’s not “your job.”
#### What You Will Work On
#### Microservices & Distributed Architecture
- Design and operate high-throughput, event-driven pipelines across a 100+ microservice ecosystem handling billions of data points
- Build and scale distributed messaging systems with RabbitMQ, backpressure management, consumer scaling, and queue health
- Develop and maintain API gateway layers with advanced routing (multi-upstream, traffic splitting, environment isolation)
- Architect SSO and identity federation for enterprise clients, supporting multi-IdP routing with zero coupling to core services
- Define clean service boundaries across ingestion, processing, and delivery pipelines spanning Ruby and Python
#### Performance & Systems
- Diagnose and resolve complex production issues (e.g., deadlocks, queue exhaustion, connection pool saturation) — and eliminate root causes
- Optimize PostgreSQL for heavy write workloads, contention management, schema design, triggers, and connection scaling
- Design and tune Elasticsearch for search, indexing, and real-time Arabic relevance at scale
- Make informed trade-offs between multi-process and async architectures based on workload characteristics
#### Observability & Reliability
- Build and maintain observability across a large-scale system using Grafana, Loki, distributed tracing, and SLOs
- Own production incidents end-to-end, tracing failures across queues, search systems, and external integrations
- Lead root cause analysis and implement preventative measures across multi-service pipelines
- Build internal tooling that improves engineering velocity, automation, deployment gating, and review enforcement
- Turn architectural principles into enforceable standards and guardrails, not just documentation
#### Platform Evolution
- Drive platform decoupling and service isolation across the system
- Contribute to Kubernetes migration and infrastructure modernization
- Standardize and improve CI/CD pipelines across services
#### Stack
Ruby on Rails · Python · PostgreSQL · Elasticsearch · Redis · RabbitMQ · Kubernetes · AWS / GCP · APISIX · Grafana + Loki
#### What We Are Looking For
- Strong foundation in distributed systems. You understand failure modes before you write the first line
- Hands-on experience with event-driven architecture and message queues in production
- Deep comfort with concurrency, backpressure, and fault tolerance
- Track record debugging complex production issues — not just fixing them, preventing them
- Experience with Rails or Python backends at meaningful scale
- You improve systems you weren’t asked to touch
#### This Role is a Strong Fit If You…
- Read and understand an existing codebase by week one
- See a broken system and fix it before anyone asks you to
- Have strong opinions about architecture and can back them up with data
- Think in systems: latency, throughput, failure modes, and cost at scale
- Treat documentation, tests, and observability as non-negotiable defaults and not afterthoughts
- Ship fast and without breaking things. Speed and quality are not a trade-off for you
- Consistently exceed expectations where meeting the bar is a floor, not a target
- Are hungry for hard challenges and actively seek problems at the edge of your limits
- Feel a sense of urgency that doesn’t require external pressure
- Have rebuilt or stabilised something significant and can talk about it concretely
#### Why Lucidya
- Real scale with billions of events and not just toy systems
- Direct impact at CTO and executive level
- Pre-IPO with clear trajectory and your work has real impact on clients