← BACK TO PROFILES

Stephen Poletto

Field CTO·Span·Colorado·

Measure the outcomes, not the tokens

Stephen Poletto

Token maxing is the newest version of counting lines of code: it is easy to pull up a dashboard and read a number, but engineers are smart, and if you reward token spend they will manufacture it in ways that do nothing for your business.

Stephen Poletto is the Field CTO at Span, a customer-facing role he describes as part marketing, part field research, and part consultative advisory, spent talking with CTOs all day about the shift to AI-augmented software development. Span connects to all the places an engineering team already works, source control, tickets, and generative AI tools, and builds a knowledge graph of what people are working on, where they spend their time, and increasingly where they spend their tokens. Poletto frames the moment with a striking image: thanks to coding agents, every organization has effectively added hundreds of engineers to its workforce overnight, an army of agents churning out code around the clock. Span exists to quantify that new activity, surfacing productivity and spend signals so leaders can see where bottlenecks are emerging.

He arrives at that vantage point after years inside the systems he now measures. Poletto studied at Brown University and built macOS and iOS applications at what he calls just the right time, riding the mobile era when those skills were hotly in demand. He spent nearly eight years at Dropbox, rising to Director of Engineering and living through the hyper-growth stretch when the company hired hundreds of engineers and could no longer rely on tribal knowledge to onboard them. He later helped scale Lattice from roughly fifteen million to a hundred and twenty-five million in annual recurring revenue. About eighteen months ago he traded a long run in San Francisco for the mountains west of Boulder, Colorado, though his CEO still lives in the city.

Poletto thinks about engineering the way an economist thinks about a market: value delivered to customers set against the cost of producing it. He argues the industry has spent decades measuring the wrong things, lines of code in the eighties and nineties, then pull request volume, each a legible number disconnected from real value. His alternative borrows from Dora, which succeeded, he notes, because it measured the system rather than the individual. So Span looks at the inputs and outputs, using LLM-as-judge evaluations of prompt clarity alongside downstream signals like code review churn and defects that escape to production. He sees engineers splitting into two archetypes: product builders who turn customer problems into requirements, and platform thinkers who, in his phrase, tune the robots that build the software. Both, he insists, matter more than ever.

What animates him now is a question he readily admits he cannot answer. He has watched AI-native teams build with a third of the people a comparable product demanded a few years ago, and he weighs that against history: the printing press turned everyone into a potential author, while robotics in auto plants simply meant fewer workers once demand for cars was met. Which pattern software follows depends, he thinks, on whether humanity's appetite for software is close to satisfied or barely tapped. His optimistic read is that talented people freed from one problem will flock to housing, health, and the harder problems society has left unsolved. Software ate the world, he says, and now AI is eating software; where that leaves everyone is the great unknown he is still turning over.

Read full transcript of interview

In this conversation: Josh Rubin (Host, CTO Studio), Jacqueline Samira (CEO, Howdy), and Stephen Poletto (Field CTO, Span).

Recorded for the CTO Studio interview series. Interview recorded on 06/29/26.

Stephen Poletto

I'm formerly a CTO and software engineering leader building software and building teams. But now I'm a field CTO, which means customer facing and helping folks navigate the transformation to AI-augmented software development.

Josh Rubin

In a tech hat?

Stephen Poletto

Yeah, I mean you get to talk to CTOs all day and I get to talk to CTOs all day. Yes, so part marketing, part field research, part consultative advisory. It's a kind of multifaceted role.

Josh Rubin

Talk to me a little bit about what Span is.

Stephen Poletto

Span connects to all the places that your software engineering team does their work. Source control, tickets, GenAI tools, and builds a knowledge graph of what people are working on, where they're spending time, and increasingly where they're spending tokens. And the idea is we help quantify productivity signals, spend signals, helping people understand where new bottlenecks are emerging, which is really valuable right now as everybody's trying to figure out the impact of AI on delivery and software development processes.

Josh Rubin

So I mean is this just productivity tracking? Are you, you know, measuring keystrokes and seeing how that corresponds to outcomes? How does this all work?

Stephen Poletto

Yeah, the most recent product that we launched is called Traces and it collects human agent interaction logs to understand agent trajectories and where they're running into issues, right? So maybe lacking documentation, lacking tools, the code base isn't really ready for agentic development. We can help analyze all those traces and identify those opportunities for improvement. So it's productivity signals, but also ideally insights that leaders can take action on.

Josh Rubin

I mean is it really as simple as this guy takes a little bit longer, he's not prompting correctly. He's his workflows are a little bit off, he's missing steps.

Stephen Poletto

So that could be part of the insight. I think a lot of the insight is also at the system level. So your code base readiness, the way that you've laid out your CLAUDE.md files or agents.md files, the way that you have kind of the environment prepared for agents. But there is actually a very strong correlation that we're finding between prompt clarity and how much you spend per line of AI generated code. So, there is an efficiency argument to individualized coaching as well. So, ideally we can provide some level of signal on both individual coaching and feedback for engineers and then system level insight into what you can invest in the code base in the environment to improve.

Josh Rubin

With everything shifting into kind of an outcome engineering-based system, is your tooling set up to essentially coach someone into being more product oriented? Like I you know you take an engineer who's been used to you know spent 10 years writing code, he's a senior guy, knows what he's looking at, has the taste for it, is now using Claude Code, Cursor, one of these tools out there. You're measuring his output and his outcomes. What metrics are you putting in place to kind of say this was a good outcome, this was a bad outcome? Or is that you're just figuring that out over time depending on the project?

Stephen Poletto

Yeah, we're the field is evolving and so we are evolving the set of data that we're collecting and the insights that we can provide. Historically in the old world it was kind of PRs shipped, you know, how many modular units of code you were able to ship and what the quality of those units of code flowed downstream. Now what we're looking at is both LLM as judge, so basically assessing the quality of the interaction, how clear is the specification, how clear is the prompting, those kinds of things. But then also looking downstream from the resulting code that's generated, is there a lot of back and forth and correction in the code review process? Do defects escape and make their way out to production? So looking at kind of downstream quality signals in addition to velocity signals. So, to your point, someone who's maybe not sufficiently business-minded or product-oriented, the way that we might identify that and help coach is to say like the specifications of what's being prompted aren't clear or you know, like there's inefficiency and a need for, you know, disengagements and human intervention in the agent trajectory and things like that.

Josh Rubin

But how are you dealing with the fact that the model might change week-to-week and something that was slightly inefficient last week is now not a problem anymore because Claude just updated something and it deals with that itself.

Stephen Poletto

Yeah, that's a great question. Like the way that we're developing these evals are to correlate them with the downstream quality signals. So, how much does it cost to produce AI-generated lines of code? What kind of code review feedback and human intervention and human time is being spent in the production of the work? What kind of defects escape out to production as a byproduct of the work? And so, we need to constantly be re-evaluating those evals alongside the model changes that are happening.

Josh Rubin

Right, because what impacts it can change week-to-week. The cost of the tokens can change. The hours that the person is working could have a direct impact on what's coming out. And you're also, I assume, using AI to determine, well, what are we going to measure this week? Or what are we going to we build in this time. So, you're using AI tooling to solve AI coding efficiency and productivity problems. And so, it's ever evolving, I imagine.

Stephen Poletto

It is definitely ever evolving. I think when I kind of zoom out and look at it from first principles, to your point about like outcome engineering and focus on outcomes, fundamentally to me, productive software engineering is that you build value to your customers at a good ratio to the cost that it takes to produce it, right? So, you've got human time being spent working on these problems, you've got tokens flowing into the code generation to bring software to life. That's sort of your cost structure. And then you've got the internal machinery of the software development process, how many PRs are created, code review, automated systems, instrumenting those is valuable to help identify what's happening under the hood. But then at the end of the day you produce software that either your customers love and the quality is high and they pay for or they don't. And so I think as we kind of evolve as an industry, we're going to see a lot more engineering productivity measurement focus on the inputs and the outputs to the system instead of just the internals, which is where a lot of the historical developer productivity metrics have focused.

Josh Rubin

Do you have some concrete examples of how the software has aided in productivity? And I ask that because I can you can very easily see how a tool like this can be turned into executive theater. It's an analytics dashboard, CYA, look how productive my org is by look at the number here. Like talk me talk me through an actual use case here.

Stephen Poletto

Yeah, so there's a couple I would highlight two powerful use cases. One is assessing all the agent trajectories across the organization. What are some common failure patterns that the agent runs into? So an example from our own internal team is that we had some, you know, lint build test commands that needed to be run in a certain way and those weren't documented cleanly to the agent. And so there were multiple, you know, interrupts, disengagements where in the agent session a human operator had to correct the agent, identify, "Hey, you need to do it with these arguments, etc." So that's kind of a system level improvement where if you make a documentation improvement or a tooling improvement, now the future agent trajectories will run more seamlessly without that kind of human investment. Another good one from one of our customers was a common code review feedback. So, you know, you're generating all this code now and all these PRs. If you have undocumented standards or expectations and then human attention is being burned on code review, if you can identify those recurring themes and say, "Hey, let's invest that into a linter. Let's invest that into an automated pre-PR gate that has to pass before the code change gets opened." Then you're saving human attention while enforcing your standards. And so, looking at those patterns of code review feedback as a way to shift left and kind of improve the scaffold for the agent.

Josh Rubin

This feels like a tool for a mature org but that isn't as agentically focused as, you know, some greenfield operation that's been up there and operating. And I think about that documentation you keep using. The documentation that you're referring to is documentation for the AI. This is not for human documentation. It's If you're documenting everything in here, so the agent reads it, you don't have to re-document this process because it's constantly reinforcing itself. Is that the right way to be thinking about this?

Stephen Poletto

Yeah, in general, things that are good for humans are good for agents. So, if there is clear documentation, if there are tools that you would expect your human operators to engage with, making sure that the agent has scope permissions to do that. But yes, I think this stuff becomes even more important because one kind of mental model that I have is you know, I was at Dropbox when Dropbox went through hypergrowth and we hired hundreds of engineers. And we couldn't invest just in kind of tribal knowledge, teaching people one-to-one about how to do things. We had to invest in systems that would create the right guardrails and the right patterns for things, you know, the easy thing to do to be the right thing to do so that hundreds of engineers could onboard more quickly and develop in a safe and consistent way. And I kind of think every organization with the power of these AI tools just got overnight hundreds of engineers added to their workforce cuz now you've got these agents that are just running and churning and generating code all the time. If they don't have the right scalable standards in place, they will misfire and you know, you'll be correcting those things with code review and with you know, quality slop and things like that downstream. So, it's yeah, the documentation, the test scaffold, basically your developer platform everything is more important than ever because you've got so much more volume flowing through the system.

Josh Rubin

Well, it's also it's an HR system for your agents. So, it like how many engineers do you have working at Span?

Stephen Poletto

We're about 15.

Josh Rubin

So, about 15 and I'm sure each of them are operating with a whole army of agents that they're working through. Is this software ultimately I mean, is that a place that it could evolve into? Like this is the software to determine are your agents running effectively? This agent is not doing what you think it should be doing. It's to give insight into your agentic workforce, which is ultimately what fuels productivity.

Stephen Poletto

If we zoom out at a high level, what we're trying to assess is the cost structure of how much it costs to bring software to life, the velocity at which you're able to bring software to life, and the quality of the resulting product. And there's a bunch of different dimensions and attributes of how you assess costs and how you assess velocity and how you assess quality, but as you start to agentify all these different stages of the process, and you're maybe trying Claude for this use case and trying Codex for that use case and building a security agent over here and a compliance agent over here, it's helpful to be able to AB test and understand

Josh Rubin

Well, not just AB test. How do I know if I have 15 agents working for me, each of them burning tokens, how do I know which one is burning the most and the most efficiently? And why is this agent so much more efficient or productive than this one? Is it something built into how this agent was constructed? Is there something it hasn't learned? It feels like it is your software is meant to kind of give insight at that level of granularity.

Stephen Poletto

Totally, yeah. Like if you have an agent with a GitHub username fixing bugs automatically for you, you could get a weekly summary of here's all the issues that this agent fixed, here's how much it cost. You know, it's kind of a report card in some sense for that agent. And so, yeah. I think that is a potential way that the product continues to evolve.

Josh Rubin

A nicer way of thinking about this is there's a reason Facebook just stepped away from monitoring all of the AI usage of their people because it sounded really creepy.

Stephen Poletto

Yeah.

Josh Rubin

Versus this is to give you insight into this massive workforce that you currently have operating underneath you. That's an interesting just from a marketing perspective.

Stephen Poletto

That's right. And I mean, if you think about like DORA, DORA was the best that the industry produced in terms of how do you think about quantifying developer productivity. And it was all about your DevOps system. You know, how fast can you get changes out? What's the reliability with which those changes go out? If something goes wrong, how fast can you recover? And I think one reason that DORA worked so well is it didn't focus on the individual, it focused on the system. And I think as we go through this next chapter of engineering productivity assessment and thinking about what it means in the AI era, we need to continue to kind of bring that system level mindset about it because it's true, individual coaching is great and that's a powerful opportunity, but at the end of the day, it's like what is my cost structure and which tools are working, which tools aren't working for producing the outcome as a team. And so, I think it as much as we can identify things like code base readiness, environment readiness, this agent's working better than that agent, you can then start to optimize at a system level.

Josh Rubin

But it also means that you're managing your system like you manage a team.

Stephen Poletto

Right.

Josh Rubin

Like you manage people. And so there are certain things that have to be considered with that. One question I'm asking everybody is you know, you've got 15 people in your org. If you had unlimited budget like in this environment and you could hire whoever you wanted to, who are you hiring for right now? What kind of person? What kind of roles? Because that's something I think a lot of people are having trouble thinking through right now since everything is changing so quickly.

Stephen Poletto

Yeah, definitely. I mean, in some ways the deep domain expertise has become less valuable than generalist abilities with some exceptional caveat that, but you know, if you think about like how difficult it used to be to program in a certain language with a certain framework to bring a certain component of product functionality online, you would have to hire somebody who like really had the depth of understanding of how those systems worked. And increasingly, if you've got good general technical reasoning skills and then business-mindedness and product-mindedness, you can get a lot done without knowing the exact syntax and the exact library version and the exact details of that domain because that's been abstracted away thanks to LLMs. So I think increasingly you're seeing folks who are like, "Hey, can I hand an ambiguous problem to this person that's a customer problem, a market problem and have them figure it out?" Right? And so that is a much more kind of generalist ability to think about the customer, think about the value, think about the business and then translate that into technical requirements, product requirements, executables for the agents to handle the details. Now I still think there's some domains that like have not been well trained in the LLM data sets where you might need that level of depth of expertise. But a lot of the kind of common, you know, application building and general software scaffolding, you need good architectural reasoning. You need to be able to think about the inputs and outputs of the system. And you need to be able to connect what you're building to the business, but a lot of the details can kind of be abstracted away.

Josh Rubin

No, it makes sense. This software is the first thing that AI is very good at doing. Everything was all built in an environment it can ingest. I mean, but ultimately what you're talking about is product.

Stephen Poletto

Yeah, I think that's really important right now. Like I think that I'm seeing engineers go down one of two career paths right now. And this is kind of always been true, but I think AI is exacerbating it. Which is you go down the product builder archetype path where you're product-minded, you're business-minded, you know how to take a customer problem and turn that into product requirements to build. And then you've got folks who are thinking about like the harness and the developer platform and the system. Almost thinking about like how do I tune the robots that build the software? Like how do I meta-engineer the system that makes sure we get quality out to production quickly. And those are more platform-oriented kind of thinkers. And I'm seeing like those two talents and skills matter more than ever. You know, people who can figure out, you know, hey, how do we make sure every change that's going through the system adheres to all of our security and architectural best practices and that we're not introducing AI slop and thinking about and obsessing about that problem, which is a little bit different than like let me go talk to customers and figure out what they need. I think you need both in this like agentic era.

Josh Rubin

One of the last questions I'm asking people is about tech debt this time around. I feel like when the AI tools started and everyone's saying, "I produced 250,000 lines of code this week." There was a big fear that tech debt was just going to run rampant and out of control. Increasingly, I'm hearing from people saying, "It's not an issue anymore. It's solved. Like, I don't have to take time away from it. Tech debt is created and destroyed in the same session half the time." what's your perspective on tech debt?

Stephen Poletto

I think agents still tend to be verbose and over-scope things, so you do have to keep them on the rails and make sure that you've got the right standards enforcement, so that you don't let that tech debt escape. But with the right investments up front, I kind of agree. Like, I've heard of a lot of innovative organizations investing in like refactoring agents or, you know, just background processes that look for tech debt and clean it up as you go. And I think like that's a very interesting pattern. The one thing that agents also are quite bad at is like knowing exactly where some reusable component exists in your stack. And so, it'll just re-author that thing. So, then you've got duplication of code and duplication of logic across your code base. So, again, like let me build an agent that specifically looks for similar functionality, similar utils across the code base, and try to refactor those and clean those up. I think like those are ways that I've seen folks make the tech debt problem fairly manageable. I mean, there is also just like a philosophical question if you if you zoom way out. Like, I've led orgs of, you know, hundreds of engineers in the past. And when you're operating at that scale, like you can't really be in the code. You look at the system in terms of the inputs and outputs, right? So, like, what's the road map that the team has promised to deliver me? Are they delivering it? What's the quality of the work? How many bugs and how many incidents and how many escalations are coming up? What's our time to respond on incidents when they occur? You know, there's a bunch of these kind of inputs and outputs to the system that you measure. And I think increasingly, there's kind of this question of if it passes the tests, it goes through user acceptance testing, it satisfies the customer, and it's a reasonable cost to build, how much does the technical like the code in the black box matter? And I think it matters in so far as it inhibits your ability to respond to customer issues or, you know, extend the system to do more to add features on in the future and so forth. So, there's still some kind of architectural primitives that I think really matter, but a lot of how people have historically talked about tech debt, it's like if there's stylistic deviation or if there's a little bit of code reuse, like how much does that matter if the inputs and outputs look good?

Josh Rubin

Last question. Is one around anxiety versus optimism? The Valley is going to do what the Valley does. It moves fast, it's building stuff. It doesn't have time to be anxious. Plenty of anxiety, but it's not around, you know, giant sociological shift necessarily. It's about can I get this thing built? Can I have my product delivered? Can I beat this other guy and scale? You're traveling a lot. The mood out in the rest of the country been nervous around all of these things. And that's going to directly impact, I think, your product because if the anxiety wins out, well, that anxiety means that your product is primarily focused on true ROI, true is this actually working? Is there value here? Does it work? And if pure optimism, if it actually, you know, golden age works out, it's about upskilling and making everyone great and going down the right path. Do you find yourself more anxious or optimistic right now?

Stephen Poletto

So, it's an interesting question because if you look at capex spend as a function of like US labor, it's quite large. I think it's somewhere between like 5 to 10%. Which basically means if you're investing in this infrastructure, you believe it's going to have that level of impact either on productivity or on, you know, supplanting the need for labor in certain places. And so, the thing that's interesting to me is like I have seen the quote-unquote AI pill, the AI native teams, and I've seen how much they're able to do compared to the software team, you know, the industry standard two-pizza team that you needed like two or three years back to build equivalent functionality. And so, to me it's very clear that for similar scope of product and software, you can build it much more efficiently. Like maybe at half to a third of the team size or something like that. Now, we haven't seen any like massive labor displacement yet. I think when you look at software engineering jobs, they're up, right? People are hiring, they're doing more. You know, you could say that's because the things that used to not clear the hurdle threshold are now clearing the hurdle threshold and you just build more than you did before. And so like that's what causes me both, I guess, some optimism and some anxiety, which is like you go on Amazon, you can buy any physical good you want for cheap, you know, it's there like and I think there's a ton of underserved software industries where like software could have an impact and it just hasn't been produced because it's been too expensive to produce. And so maybe we'll just have like a ton of new software emerge that solves problems for us in our daily lives. But how much of that is really there? And as teams transform and truly become AI native, what is the real math on the labor displacement? I think like that's the big unknown right now. And so the optimism is like, well, we can build a lot of stuff we never thought we could build and maybe a lot of talented people will surge toward problems that we should be solving as a society, but the negative is and a lot of this stuff can be done with way fewer people. So like what's actually going to play out in practice?

Josh Rubin

Yeah, third and fourth order effects are

Stephen Poletto

They're really hard to predict right now, you know? And I think everyone is looking at kind of the jobs numbers right now, but that's a present indicator, not a future indicator of where we'll be. If you talk to any of these truly AI native teams, it's like the way they bring software to life is so different in terms of staffing than a couple years ago, and I don't think we've seen those effects like really play out at scale at all yet.

Josh Rubin

Jacqueline, do you have any questions? Are you still in here?

Jacqueline Samira

Yeah, I'm here. I want to pick up where you left off where you said, you know, now we have so much capacity and engineers are flying through and they can build so much. The problem is though that means we need less people to build it.

Stephen Poletto

Yeah.

Jacqueline Samira

And the question I have for you is like if you go back into the past and you look at every single revolution that's come before, usually more is more. And so like when they built the printing press everyone like freaked out and they said, "Oh my gosh, you know, we're losing so many jobs." But then what ended up happening is now everyone could be an author because publishing became so affordable. Or more affordable than it had been when things were handwritten. And so my question to you is when you look to the future if there if that is the reality, if it does explode and now all of a sudden yes, we don't need as many people because they're producing more, but then our ideation of what can be done becomes greater.

Stephen Poletto

Yeah.

Jacqueline Samira

Where do you see and I and I heard you talk about like the bifurcation of the product business oriented engineer and then the platform engineer. But like living in this third order effect that I know is very opaque for all of us, where do you imagine it goes and what type of traits if people are going to school right now, if they're going and getting an education right now, that would be valuable in this future in your opinion?

Stephen Poletto

I do not have a crystal ball. I think these are like the super interesting questions. Like I do think in many ways like the progression of the software industry has been a progression of democratizing the ability to create software cuz like we used to do punch cards and then assembly and then there were higher level programming languages and then you know, people have built application frameworks. Like you want to build an iOS app, Apple has provided all of these toolkits out of the box. So it's really easy. And as things have gotten easier and easier and easier, there's only been job growth and job creation. So, I think like that's the looking backward, that's the optimism looking forward that as we continue to make it easier and there's platforms that allow you to basically program in English language to find what you want to build and you bring it to life that people just create more stuff. And more stuff makes all of our lives better. And I think that is the optimistic framing. I think the question for me is like what is our collective need for software? Like where does What is the satiation point where we've fulfilled the demand? Right? And I think you cited like the printing press and how it democratized all this creation. You know, the alternative would be to look at like the automotive industry where you had robotics introduced into the plants creating the cars. And the role of people working in the plants is now to tune the robots. And it's kind of meta engineering to create the system where actually through automation and robotics the cars get created. And that could be a kind of apt metaphor for what's happening with software engineering. But there's only so much demand for cars. Right? And so, what happened was a lot of people lost their jobs because there was only a certain you know, demand for how many vehicles are needed by the world. So, I think the big unknown is like we and we I don't think we've seen the end. Like software is eating the world, now AI is eating software. Like it's possible that there's just demand for software for the foreseeable future and we live in a very optimistic world. But at what point are we kind of like, okay, like the important problems with software have been solved? And I could even have an optimistic framing to that, which is like, okay, then a bunch of talented smart people go work on other problems that matter, you know, like we've still got a lot of health issues to work on, housing is a problem. There's all kinds of societal problems that smart people could flock to. So, you know, there might be some short-term disruption, but then maybe in the long term it just means talent flows to the areas of most need.

Jacqueline Samira

Do you find that your team is excited about what's happening right now? Do you find that there's any fear? What's the general sentiment of your team and how has their feelings changed the way you've had to lead?

Stephen Poletto

I think it's a little bit of both. There's a little bit of excitement and then there's a little bit of recognition that what people talk about on Twitter and in the headlines is with an incentive to sell something and overblown and kind of like some healthy dose of skepticism about the reality. But I do think our team is like embracing the change and rolling with it. A lot of the organizations that we work with have had their engineering teams opt in the same way. Like they're curious. Yeah, there's some anxiety, but they're curious and they're willing to learn like a new way of working. You asked me earlier about like what advice I would have for folks like studying in university. And I think one is critical thinking. It's always been important, but now with how easy it is to create content and misinformation, like the ability to reason critically and get to truth matters more than ever. But then I would say like embracing this new way of working and treating it as a tool, right? Like I benefited in my career from working on Mac OS apps and iOS apps at just the right time when like the mobile era was taking off and it meant that my skills with building iOS applications were hotly in demand and like I benefited from that personally. I think right now AI skills and the ability to use these tools to their full potential enables you to be way more productive, enables you to generate way more, and employers are looking for that. So, I don't think that the way through the anxiety is to be afraid of it. I think it's to lean into it and figure out what this means first hand.

Jacqueline Samira

Last question. What are your thoughts on companies that are token maxing?

Stephen Poletto

Token maxing is dumb.

Jacqueline Samira

Please tell me more.

Stephen Poletto

Okay, so I believe that the companies in the headlines that have optimized for the most legible metrics are taking the easy road. It's the same old story when folks tried to quantify developer productivity through the '80s and '90s and 2000s, they focused on lines of code, which generated verbose code that nobody wanted to maintain. And then they focused on, you know, pull requests and volume without a connection to the value that those changes are producing. And I think token maxing is basically the newest version of that. It's easy to measure. You just pull up a dashboard and you got a number. It shows how like innovative you are, whether you're like you know, ahead of the curve or behind, you know, oh well, we're spending a lot on AI, so we're ahead, you know. But fundamentally, it creates us all these perverse incentives. I do not think that the most efficient way to use these AI systems is to max your tokens. There's ways to make your token utilization efficient. And we've heard like kind of the stories of people at some of these organizations who are gaming the number. You know, Goodhart's law says that when a measure becomes a target, it ceases to be a good measure. If you've got token maxing leaderboards, engineers are smart. They're going to figure out how to manufacture token spend in a way that is not useful to your business. So, I think it's a terrible idea. I think it was designed to promote adoption and be a blunt instrument to get people using the tools, but I think the shelf life on it is going to be really short because of all of these like second order effects that it creates.

GET INVOLVED

Be part of the
conversation.

Whether you're a CTO who wants to be featured, a company looking to sponsor, or an engineering leader wanting a seat in the room — there's a place for you here.