Job Description
Company: Day One Partners
Location: New York metropolitan area, US
We’re partnering with an early-stage, venture-backed AI startup building the collaboration layer for humans and AI agents.
The company is developing systems that allow agents to maintain context, work across long-running tasks, collaborate with people and other agents, and become more effective over time. This is a small, highly technical team tackling problems at the edge of what production agent systems can reliably do today.
They’re hiring 1–2 Applied AI Engineers to take significant ownership of the intelligence and infrastructure behind the product.
About the Role
You’ll build the systems surrounding frontier models that determine how agents remember, act, coordinate, and improve.
The core scope spans agent memory, task optimization, agent orchestration, harnesses, evals, and benchmarks, alongside the infrastructure required to run these systems securely and reliably at scale.
This is a hands-on engineering role for someone who wants to work below the application layer. You should be excited by the hard parts of making agents actually work in production, not just integrating an LLM into an existing product.
What You’ll Do
– Design and improve agent memory, context management, task execution, and multi-step orchestration
– Build and optimize agent harnesses, including tool use, prompting, planning, feedback loops, and failure recovery
– Develop evals and benchmarks that rigorously measure agent quality, regressions, and improvements
– Build infrastructure for running and scaling sandboxed agents, including isolation, observability, reliability, and cost/performance
– Own difficult technical problems end-to-end, from experimentation and architecture through production deployment
What We’re Looking For
– Demonstrated experience building agent systems, harnesses, evals, benchmarks, or closely related AI infrastructure
– Strong software engineering fundamentals with depth in backend, infrastructure, distributed systems, or production AI systems
– Ability to reason rigorously about agent performance: define success, design experiments, diagnose failures, and prove whether a change actually improved the system
– Experience thinking through scalability, security, reliability, and isolation for complex production workloads
– High technical ownership and comfort operating on problems where established abstractions or best practices may not exist yet
Strong Signals
We care considerably more about the depth and quality of what you’ve built than a particular number of years of experience.
The strongest candidates will have done things like:
– Built agents or agent infrastructure that real users depend on in production
– Designed an evaluation harness or benchmark and used it to drive measurable improvements
– Worked deeply on memory, orchestration, tool use, long-running execution, or agent reliability rather than only prompt/API integration
– Built technically demanding systems where latency, concurrency, security, reliability, or scale meaningfully affected the architecture
– Taken ambiguous technical problems from first principles through experimentation and into production
Why This Role
You’ll work on problems where the underlying models are improving rapidly, but the infrastructure around them is still being invented.
How should an agent decide what to remember? How do you optimize performance across a task that may run for hours? How do you benchmark something nondeterministic? How do you safely run large numbers of agents capable of taking real actions?
If you’ve already spent meaningful time wrestling with problems like these and want substantially more ownership over them, we’d love to hear from you.
Source: LinkedIn