New Relic remote (Portland OR based) Lead Software Engineer for Data Platform ingest into NRDB, the telemetry datastore handling metrics/events/logs/traces at massive scale. Java-first Kubernetes fleet with Kafka backbone and cell-based architecture; millions of messages per second, petabytes of data. 7-10 yrs exp incl. 2+ yrs lead; Kafka, Kubernetes, Argo CD, relational query optimization. No visa sponsorship; export compliance assessment may apply.
Requirements
7-10 years of professional software engineering experience, including 2+ years leading or acting as a technical lead for a team or major initiative
Deep fluency in Java or a comparable OOP language
Experience running stateful distributed services in production under unpredictable load, including debugging complex failure states and captaining incident remediation
Experience leading architecture and design reviews for complex systems
Hands-on experience with cloud infrastructure, CI/CD tooling (Argo CD or similar), event-driven architectures (Kafka or similar), and container orchestration (Kubernetes or similar)
Experience with data modeling and query optimization against relational databases at scale
Track record of mentoring engineers
Strong written and verbal communication to non-technical stakeholders
Visa sponsorship not available for this position
Export compliance assessment may be required for certain roles (encryption software subject to U.S. export controls)
What you'll do
Set technical direction for New Relic's Data Platform ingest team: own architecture and design decisions, weigh tradeoffs, lead engineers through complex problems in high-throughput distributed systems
Scope, estimate, and own delivery of complex multi-engineer projects from design to GA
Evaluate and introduce new tools, patterns, and technologies including LLM-assisted development practices
Stay hands-on: write and ship production code alongside the team
Develop engineers through code review, technical feedback, and career mentorship
Communicate technical risk and tradeoffs in terms stakeholders can act on
Own operational health: CI/CD, observability, incident response
Close gaps in tools, documentation, and team coverage