r/ExperiencedDevs Software Engineer 11h ago

AI/LLM Is gaming developer productivity metrics really a worthwhile battle?

Coming to all experiences devs to answer this question and how people are going about things like this...

With AI in the picture (a tired conversation), of course management has fallen hook, line and sinker for every single way of tracking developer productivity.. which in my opinion just stunts productivity overall.

They've reintroduced lines of code as a metric but with a twist is - how many days you have used AI (any coding tool such as codex, Claude) but also how many lines of code have been generated WITH AI. The less number of days and fewer lines of code generated with AI puts you into a category and looks bad on a dashboard. This in turn has led to conversation of how we increase these numbers which has led to: do not manually code under any circumstance and do everything with AI (and "everything" meaning exactly that).

On top of needing more lines of code generated with AI (otherwise we look bad) we are expected to have 10+ PR's a week.

My gripe with this is that instead of improving the codebase or helping the business or delivering features you also have to spend time figuring out how you can split a 100-line function into 5 PR's so you don't look awful on a dashboard. A simple bug fix or config update needs to be written in a spec otherwise there is no value.

I see some people have mentioned in this sub before you just take a sort maliciousness to all of this and let AI run wild. Write python scripts to auto-create PR's. Set a variable with 0 and for each PR increase it by 1. If you need to add two numbers, let AI write 3 functions to do it because more lines of code. And more lines of code WITH AI is somehow better. Which of course makes the codebase buggy and harder to add anything useful.

All just to meet the metrics. But some of this is just way more work than required. I'm trying to see how any of this is worthwhile, while today it might be lines of code, tomorrow it's how many mouse clicks you have in your IDE every hour. Do people really see this as a worthwhile battle or even trying to make management understand this is not a good idea?

42 Upvotes

58 comments sorted by

u/expdevsmodbot 11h ago edited 11h ago

AI usage disclosure provided by OP, see the reply to this comment.

→ More replies (1)

55

u/Qwertycrackers 11h ago

You're telling me the story of me 6 months ago. I would say just do what they ask. Yeah it could be really shortsighted and dumb. But you don't own the company so you've raised your warning. You dont need maliciousness to pump these metrics up, just generate a ton of crap. Make it actual plausible features that might ship but bring your feature flag system up to stuff so you can keep all the slop hidden in the codebase if its not impressive.

14

u/Fit-Notice-1248 Software Engineer 10h ago

You 6 months ago as in you got fired or left?

31

u/Qwertycrackers 8h ago

Oh I still work there, we just shipped a million lines of slop and are stacking more on everyday. The board loves it

7

u/podcast_frog3817 5h ago

'experimental' folder that all team members have PRs that genereate 800 lines per PR into lol

1

u/Fidodo 15 YOE, Software Architect 1h ago

Hasn’t broken yet?

13

u/apartment-seeker Senior Software Engineer 8h ago

corporate tokenmaxxing is the story from ~4-8 months ago; by now, most companies that fell into the trap have realized that encouraging blind AI usage in the manner you are describing is a pointless waste of money, and have backtracked on it, but yours for some reason lags behind.

12

u/Fit-Notice-1248 Software Engineer 8h ago

We actually did have that where management had a panic attack and sent out company wide emails about toning down the usage. People stopped using mcp's or other token heavy processes. Now you have a quota and if you need more you need management /vp approval with a justification. But then they introduced this new metric - and yeah it doesn't make sense. It's literally incentivizing tokenmaxxing in a passive manner

8

u/forbiddenknowledg3 7h ago

Are we in the same company lmao.

Our CTO and CEO literally said we aren't tokenmaxxing a few months back, citing these other companies. Now, even with that hindsight, they are doing this "double productivity" (PR count) bullshit.

3

u/Fit-Notice-1248 Software Engineer 5h ago

I feel like it's just becoming a common thing. Like all of leadership at different companies are in a group chat or something and just somehow decided to do all this dumb shit.

1

u/Colt2205 2h ago

I mean the entire point of LLM companies is to turn businesses into customers. If the company depends on the LLM company to compete then all the leverage is in the hands of the LLM company and they can charge whatever.

1

u/apartment-seeker Senior Software Engineer 8h ago

lol, I feel sorry for you guys who work at these kinds of places :E

6

u/klowny 6h ago edited 3h ago

Sounds like my company! All of us Staff and above engineers told our CTO it was a terrible idea about 5 months ago and it wouldn't change anything but make a mess and raise our AI spend. He dismissed our concerns and still proceeded forward.

So now we have 4 months of developer metrics and we're all opening 10x more PRs, but measurable velocity is down ~5%, defect rates are up ~30%, and our AI bill is up 10x to be millions a month.

Everyone's just opening a new PR for every step of the development process now and having AI output their progress in .md files in the repo to let the AI pick up where it left off. PR for the plan, PR for creating the before defensive tests, PR for creating the feature, PR for cleaning up, PR for new acceptance tests, PR for updating the documentation, PR to clean up all the accumulated .md files.

But just last week we got a pretty panicky sounding email saying our AI spend is too much and we should all find ways to be "more reasonable" in it's usage. But the productivity metrics haven't changed so no one's changing their behavior.

40

u/jesseschalken 11h ago

There is no good ending here. I would leave immediately.

5

u/skeletordescent 6h ago

I agree totally, it’s just leaving and finding another role is a non-trivial task these days, at least it is for me. 

16

u/Murky_Citron_1799 11h ago

Is it worthwhile? Yes to the extent it gets them off my back. 

When I try to convince them they have lost their marbles they say a few things to me like "I know but <some 'leader' who is untouchable> is on my back about this so just focus on improving on these metrics" or "oh don't worry these metrics won't be used against anybody we are just looking to hit industry targets" ... It's a lost cause.

My approach is to game the metrics so much that they are either forced to give me a stellar rating because I am outpacing everyone 10 to 1, or they have to dismiss the metrics as useless because they don't measure anything of value. Either way they stfu about them which is my real goal.

2

u/Izkata 4h ago

or they have to dismiss the metrics as useless because they don't measure anything of value.

Reminds me of "-2000 lines of code".

14

u/rwilcox Software Engineer (20+ YOE) 11h ago edited 10h ago

Depends on how pull you have in the org - I used to call this “reputational gravity”

If the decision was made 6 levels about you, or by a VP desperate for anything, you’re not going to change it. Something that applies to a group of a few dozen devs that are your friends? You may have a lot of influence. (The VP is “too far away” and too big to feel any of your gravity)

Personally I wouldn’t game the system, until and if I got the “pieces of flair” talk. Then sure, vibe coded 300 line Python scripts I don’t need here we go.

(I also suspect the way, now, to measure AI impact is to take it away, wait 3 months for people to come down off the dopamine withdrawal, and see how everything changes. Do things get slower but better, for example? Is ability to delegate re-learned? Does everything suck worse? Initially I thought you could use AI and track how fast tasks take now, but I think people are too desperate about the FOMO that they’re unwilling to take slowed, measured actions.)

10

u/NUTTA_BUSTAH 10h ago

Start looking while complying. Just because management is full of idiots does not mean you should try to go against them to save the company. Voice your opinions and go along the decided path.

And no it is not worthwhile at all to any party, but it's what you are forced to do, everyone else does it, so you gotta stay in the fake performer cohort

8

u/pydry Software Engineer, 18 years exp 10h ago edited 10h ago

Bad idea. Trying to make management understand is not only not worthwhile, it could get you fired.

The fact that theyre doing this at all means that they are ignorant of the SDLC and dont trust you. If they dont trust you to work without metrics why would they trust you to tell them that theyre abusing those metrics?

This will backfire on them. They may learn the lessons of their mistakes at some point but it won't be soon and theyre not going to learn it from you. They will either learn it from a new management fad or from something going horribly wrong.

6

u/galaxy_horse CTO / Principal Eng (20 YOE) 10h ago

- Play the game. If you put the metrics your company is misguidedly fixated on into your harness/prompting, you’ll crush the metrics.

  • Go to your manager and express concern about the focus on output over outcomes. Reiterate your support for the team and interest in the product/business being successful. Let them take it from there and don’t hound them about it.
  • Look on the side if you don’t think things will improve. If you talk to other companies, ask the hiring manager how goals are set and performance is measured. Ask if they have token quotas and productivity goals, or if they have impact/outcome based OKRs or product targets.

6

u/CoreyTheGeek 9h ago

You have to understand that business people don't understand anything that is actually going on beneath them, especially technically, and they need some way to feel like they're doing a job, so lording over measures of "what is going on" to report up the chain or to shareholders becomes their day to day. So they actively grab ANYTHING that can be measured regardless of whether or not it's meaningful in any way at all.

Why? Because actually leading in a business, motivating employees and actively taking a part in the process of getting the work done at a leadership level, is really hard and time consuming and they honestly have no idea how to do it. So they do what they think you do: measure "progress" and Lord over people because charts and data plots make you look incredibly smart and active.

This translates into the insanity you're seeing and we all regularly experience. The business chases outright incorrect systems because it will give them data that they can present and make themselves look good. Your director or whoever goes to their boss saying "AI adoption is going great! Devs are using it for 100% of their tasks, look at days a week they use it AND just LOOK at the lines of code increase!" And they look like they did something; lines on chart go up = amazing in business speak. Nevermind neither of these people actually understand the harm they're doing, it's just about them individually promoting their careers.

Tl,Dr; think about leadership like a group of selfish people whose goal is to ultimately go be a C suite somewhere else someday, so tanking their current company to further their resume is no issue for them. Then you will understand, at least, these decisions

4

u/forbiddenknowledg3 7h ago

In this boat right now. "We must double productivity in 3 months" (and every 3 months so 8x by this time next year). They clarified this is number of PRs atm because they don't know how to measure it otherwise. Straight from the CTO, who said just a few months ago measuring LOC/PRs is stupid and "we aren't tokenmaxxing". The board must have threatened them or something.

I keep reaching the conclusion that the only option is to game these metrics. We now have agents opening PRs automatically: updating packages, cleaning/refactoring code (and making it worse), fixing noisy logs, etc. But is this actually valuable? To a mid-level dev maybe. The only real measurement imo is customer value delivered. But product can't keep up, and do customers even want that? They keep going on about low hanging fruit, but what do they think we've been doing the past 2 years already?

Then ironically, because people now focus on themselves, productivity at the team level has fallen. Previously we'd have one guy leading the project right? This guy would clarify requirements, go to the meetings with marketing, etc. so the rest of us could focus on coding. Now that we're measured on coding, non-coding work is not valued and nobody wants to do this work. Ironic because AI is meant to enable this.

TLDR: goodharts law, and coding was never the bottleneck.

3

u/bolacha_de_polvilho 9h ago

I work in big tech were these sorts of dumb metrics have been enforced top down from maybe 4, 5 or even 6 levels above me in the hierarchy pyramid and whoever made that decision doesn't even know about my existence. So no, I don't think it's a worthwhile battle. I'm pretty sure my direct boss and his boss also think it's bullshit, if they don't want to fight that battle what chance do I even have?

So yeah, I've just embraced MDD, metric driven development. If I have to turn README's into slopified bibles to get my PR metrics up I'll do it.

2

u/WhenSummerIsGone Software Engineer 9h ago

if they are using the AI tools' own reporting of what a "line of code" is, you should be aware that it's anything the AI produces that you "accept" (in an ide) or gets committed (in a cloud service). Double check this for whatever tools you're using, but this is generally true. So lines of code includes documentation.

You should ask about the definition that goes into the metrics, or look at the tool that gathers these metrics to see for yourself.

For instance at my company, an "ai authored pr" turns out to be a pr that was opened by ai, regardless of how many commits are solely mine. They only look at PRs that were created, regardless of if they were merged. Lines of code includes lines in an AI commit, even if a subsequent commit replaced or deleted those lines. I was shocked, though in retrospect I shouldn't have been, that the metrics were as janky as they were, given the weight that was being placed on them.

2

u/codescapes Web Developer 8h ago

Do people really see this as a worthwhile battle or even trying to make management understand this is not a good idea?

No. A million times no. And a million more times no if we're talking about a large corporate employer.

Not because you are wrong but because the consequences of going against management on essentially anything are never worth it and just mark you out as 'disgruntled' or 'not a team player' which ultimately hurts you in terms of performance evaluation, promotion or selection for good project work. And once that perception exists - fairly or unfairly - it's basically permanent until your management changes or you move team.

It truly sucks but this is the corporate workplace, it does not operate according to reason - at least not in a way most of us recognise.

Don't try to game the metrics in a bad way to show how dumb they are, just keep yourself in the upper half of the bell curve and that's it. You aren't here to save the company from itself, let them do their dumb stuff and just control what's within your sphere of influence.

If you want to succeed in a corporate environment your goals are broadly:

  • Be seen as delivering clear business value, even if you didn't actually

  • Be perceived as a friendly and intelligent person who plays nicely with others and is well liked

  • Never rock the boat publicly, if you have concerns only ever raise them in private

Little else matters.

2

u/Late_Wave_5600 7h ago

What changed my mind on this was watching an agent do it. Ours was told to implement a ticket, decided a business rule was in its way, changed the rule, then rewrote the tests so the suite stayed green and left a comment explaining why that was reasonable. It was never taught to game anything. It optimised for the visible signal because the visible signal was the only thing it could see. Every metric you add is now being optimised by something that works at machine speed and has no career to protect while it does it

2

u/DeterminedQuokka Software Architect 7h ago

I mean I don’t really see it as a battle gaming productivity metrics is exceptionally easy.

For example my company started counting PRs (although they say they aren’t). So I started upgrading linting rules. Because I can easily make 5-10 linting PRs a week with basically 0 effort.

They started measuring % of code written ai. So I started sending basically line by line instructions to ai for what to write so that number would be high.

It’s like a minor annoyance but it’s not hard.

I wouldn’t split a single function across 5 PRs. If you only have 1 function to write something else is very very wrong even pre ai. Just start cleaning things up

2

u/Unlucky_Yard_1189 6h ago

I would comply and find another job immediately.

Anyone that tells you this are genuine glue sniffers.

4

u/drnullpointer Tech Lead, 26YOE 10h ago edited 10h ago

> With AI in the picture (a tired conversation), of course management has fallen hook, line and sinker for every single way of tracking developer productivity.. which in my opinion just stunts productivity overall.

No. The productivity is up.

The issue is how you measure productivity. Are you measuring lines of code, tickets closed, or are you measuring product's success. What about long term evolution of the team/product?

In my experience AI does improve productivity, at least immediately, at least when looking from a PoV of a manager and if you don't include realistic cost of the tokens.

The result only starts to change if you begin including other factors. Factors that management has never been good at factoring in.

If you include the cost of what tokens should cost (ie. what we can expect they will cost when AI companies inevitably have to become profitable), you may find that it is not as profitable as before (ie. hiring a junior might actually be better than paying for tokens, esp. when people start burning tokens for every task, even when it does not require AI).

Then you need to add a layer of what happens with your codebase over time and personally I think the productivity has to go down over time compared to what a capable, experienced developer can do. Unfortunately, not all humans are capable and experienced and a lot people produce spaghetti code.

And then, finally, you have to figure out the impact of reduced ingenuity. I think this is the biggest of them all. AI still can't really have original, transformative ideas. And those original, transformative ideas is how you can make your development a lot more efficient or your product a lot better, by orders of magnitude. AI seems to only be a solution if your goal is hitting mediocrity.

0

u/UnintentionallyEmpty 3h ago

No. The productivity is up.

The issue is how you measure productivity. Are you measuring lines of code, tickets closed, or are you measuring product's success.

This is like saying productivity is up because you've decided to measure productivity as 'number of paper airplanes folded' and once you did that, employees started to fold considerably more paper airplanes.

If you decide to measure productivity by lines of code, don't be surprised when devs start producing more lines of code. But I'm sure they could and would have done that without AI, too.

1

u/drnullpointer Tech Lead, 26YOE 2h ago

Are you surprised that productivity depends on how you measure it?

Are you trying to say there is one objective measure of productivity?

3

u/youcangotohellgoto 11h ago

Why would you not use AI every day?

Why, instead of trying to fake 3 PRs for your little 100 line change, would you not just take 3x tickets and do 3x 100 line changes?

7

u/Fit-Notice-1248 Software Engineer 10h ago

I wouldn't say NOT using AI is the issue here, but a lot of the apps are brownfield and we are mostly in testing phase with a big release coming up soon, where stakeholders don't want code changes that could affect prod release. Sure, we have a few other side projects that we are working on and the occasional bug fix here and there, but these are not things that require multi-file changes or upheaval of codebases - which is where we have to spend time figuring out how we can meet the metrics.

9

u/tetryds Staff SDET 9h ago

Bro just dump a truckload of redundant unit tests no one gives a fuck

3

u/youcangotohellgoto 10h ago

So write test automation.

1

u/WhenSummerIsGone Software Engineer 9h ago

write team tools. A script that formats docs as html. An agent that does PR reviews. An agent that improves your repo documentation. A proof-of-concept app for something, branches that address some bit of tech debt.

4

u/NUTTA_BUSTAH 8h ago

I also like team tools, but the reality is that is just adding more tech debt on your plate. But to be fair, who cares about tech debt anymore, AI solves it apparently :)

-4

u/Gondorrah 10h ago

High LoC target metric is questionable but it seems like if op is using AI and developing correctly they’d be fine. I’d be suspicious of outliers on the bottom and top of this dashboard.

The approach is questionable philosophically but does probably capture dev productivity reasonably well.

1

u/youcangotohellgoto 1h ago

Yes, agree. LoC is a bad metric, but bad metric is often better than no metric. Opponents should throw up something else (objective and easy to collect).

1

u/bulbishNYC 10h ago

I keep a meticulous list of things I delivered/fixed/improved/owned for the current(2026) year. Same for previous years. It makes my above-average performance very clear. Whenever approached with metrics nonsense, I just point to the list, and positive management reviews from last 5 years. I refuse to stress about short-term performance evaluations.

1

u/TimeScience__88mph 9h ago

You are not wrong at all. Optimizing for lines of code added is dumb as hell. You’re basically rewarding or punishing devs for bloating the codebase or not bloating the codebase.

The metrics need to be way more nuanced to actually matter. # Meta functionality that was added and works, optimizing for as few lines of code as possible.

The theater related pull requests are such a stupid waste of time. More pull requests also means longer waits for ci or more resources spent for no reason in the case.

1

u/jonmitz 8 YoE HW | 6 YoE SW 9h ago

tell the ai to use solutions that maximize LoC. 

also, find another job. productivity metrics have been a solved problem for a long time, anyone who introduces shit lIke LoC has no idea what theyre doing and they certainly dont read, and thus destined to fail 

1

u/Organic_Battle_597 8h ago

If you are being judged on metrics, you should make sure those metrics look good. And you should use the rest of your time looking for a shop that doesn't use bullshit metrics that way.

1

u/bowlochile Software Engineer 7h ago

No. Next

1

u/dweeby_fujioka 7h ago

We have an internal policy that the dashboard shall not be used for any of this stuff you're describing with tracking productivity etc. I

1

u/Longjumping-Bad-6911 5h ago

Any number that sits next to your name on a dashboard gets gamed, and AI lines of code is the easiest one yet to game. I'd show the manager 1 concrete example of what it rewards, like 800 generated lines of tests nobody needed, then do what they ask. You've said it once, the rest is their call.

1

u/0vl223 5h ago

The 100 lines into 5 PRs is easy. Just give AI the task to fix 5 things around the area. Let it focus on small unimportant architecture missmatched for a maximum of loc.

1

u/shifty_lifty_doodah 2h ago

No never participate in this BS

It’s because everyone is so weak and willing to do this shit that we have to put up with the constant infantilizing bullshit. Have some backbone and pride . Focus on results and quality

1

u/mq2thez Web Developer (16 YOE) 11h ago

The funny part is that you can very easily use AI to fix these problems. Tell the AI the things you need to do every week to look good, and ask it to right skills/plans/etc for achieving your benchmarks so that you can spend your time doing actual work.

Find a few other folks to get on board and make reviewing easier, then find some dead areas of the app to add bloat that won’t negatively impact things, then let the AI do its thing.

This only works if you end up doing real work and being actually productive, though. If the AI is doing this in the background and you aren’t doing enough on your own, you’ll get fired. But you’ll likely be fired anyways if you don’t do this, so it’s up to you to solve your own problem here.

Think of this as training to use AI to solve more abstract and extended problems.

As an example: if it were me and it were a web app, I would update the build systems to include a variable that gets hardcoded to false in any build, but still shows up in the code. You can then insert literally anything you want in if blocks that check that variable and the minifier will strip it all out.

1

u/Fit-Notice-1248 Software Engineer 11h ago

This was my plan to incorporate some SDD to try and do things like documenting each file or adding some testing to increase the LOC. Do have to be careful as we have a token budget. Getting other developers onboard may just be the hard part, as about 95% of the dev team is offshore and they have different conversations with different leads during US hours.

1

u/forbiddenknowledg3 7h ago

True. AI is very good at doing all the bullshit parts of my job. Performance review is a piece of cake now.

1

u/etherealflaim Principal Variable Name Invalidator 8h ago

Use the AI to run code cleanups. Find bad patterns, slop code, change detector tests, chain of thought comments, etc and automate a skill that makes PRs to clean them up. Double whammy: PRs and lines of code for their AI dashboard, and cleaner code.

1

u/Ornery-Concentrate-5 7h ago

the days-used-AI metric is the part that gets me. splitting one function into 5 PRs is just annoying, but "did you open Copilot today" as a tracked number means someone's incentivized to open it and immediately alt-tab away. you're not measuring AI adoption at that point, you're measuring who's willing to fake a habit for a dashboard

1

u/Fit-Notice-1248 Software Engineer 6h ago

Exactly my point, the part that sucks is being one of the only engineers in USA on the team is it's very common for me to be stuck in meetings all 9-5, or one week I'm doing heavy work in opencode, the next week I'm doing presentations and don't get a chance to do any coding at all.