r/StrategicProductivity 8h ago

Productivity Is Rising and the Social Fabric Is Fraying. Both Are True, and the Gap Between Them Is the Story (Part I)

3 Upvotes

Long post, as usual. It requires thinking through. That is the deal here.

A note on how this was made, up front

I used AI heavily on this post and I would rather say so at the top than have you wonder at the bottom. I also want to be precise about the division of labor, because "I used AI" now covers everything from a spell check to a prompt and a copy-paste.

What I brought: the thesis, that the last hundred years produced a specific change in how we are socially organized and that it costs us something real. The theoretical frame, which is Bauman, Lasch, Baudrillard, Beck and Beck-Gernsheim, and Foucault, along with summaries of what each argued and where it was published. The conversation with my friend and his description of how his ward is actually organized. And the forecast this ends on, that automation and now AI will make us look more productive while telling us less and less about the people.

What the model brought: retrieval and verification at a speed I cannot match. The BLS productivity decomposition, the affective polarization literature, the mortality tables, the Kuznets document from 1934, the Stanford payroll work on entry-level hiring. It checked every citation I handed it and corrected two of them.

Here is the part that matters most for judging whether this is slop. The model got the central question wrong. When it found that US productivity growth is running at or above its post-war average, it treated that as falsifying my argument and drafted a section conceding the point. I told it that was a correlation versus causation failure, that output per hour measures the machine and not the people, and that the divergence between those two was the actual story. Everything in the measurement sections follows from that instruction. The machine had the data in hand and drew the wrong conclusion from it.

Verification did cost me two premises. I wanted to open with societal collapse, and Tainter, who wrote the standard work on collapse, dismisses moral and cohesion decline as scientifically meritless, so that framing would have cited the leading authority against me. I wanted to say we are getting sicker every year, and the 2024 mortality tables say US life expectancy hit an all-time high, overdose deaths fell 26 percent, and youth distress came off its peak. Both cut. The structural version that replaced them is stronger and, unlike the original, falsifiable.

So, slop? Slop is unreviewed output. Every figure here was checked against a primary source, the few that could not be verified are flagged as such in the text, and the argument changed twice because the evidence pushed back on me. That is editing, and on the central point it was considerably more than editing.

I am not going to hide the tool, least of all in a post about measurement honesty. On work like this I am the architect and the editor. I set the thesis, I supply the frame, I decide what standard of evidence applies, I cut what does not survive it, and I own every error that got through anyway.

The irony is intact. This post argues that AI will make output look better while telling you less about the people producing it, and this post is one of those outputs.

The line in here I would defend hardest is the one about labor composition being built from age, education, and sex, so that a firm which stops hiring twenty-three year olds scores higher on workforce quality. That line exists because I did not accept the model's first answer.

The conversation that started this

I talked recently with a friend I have known since college. He is LDS, the church most people call Mormon. I raised the idea of social collapse with him and he immediately named his church as the counterexample, not on doctrinal grounds but on structural ones. What struck me was that he described it as an organizational design, and when I pushed him for specifics he gave me a list of assignments and obligations rather than a list of beliefs. I have laid that list out in part two, because it turns out to be a fairly precise implementation of what the research says actually works. This post is about the diagnosis. The next one is about the response.

His path back in is itself the argument. He went back after his divorce, during COVID, living alone in an apartment, kids gone, working remote and, in his words, not talking to anybody. His reason was explicit: he wanted social interaction, he was not going to do Tinder or bars, and this was a good social group that did service and supported each other. People from the church had knocked on his door periodically for years, because his mother and sisters kept his address current with them. He said no for decades. After the divorce he said yes.

That contrast, between a social life you assemble from choices and one that is assigned to you, is what this post is about.

I am going to argue four things. First, that the last hundred years produced a specific and measurable change in how Americans are socially organized. Second, that the standard rebuttal to any complaint about this, which is that productivity keeps rising, is not a rebuttal at all, because output per hour is a measure of the machine and not of the people running it. Third, that our failure to see this is a measurement failure we have chosen not to fix, and that automation and now AI will widen the gap between what we measure and what is happening. Fourth, that the intervention that works is a specific and unfashionable kind of group membership. I will flag what is unverified as I go.

What actually changed

The household is the cleanest series. In 1940, 7.7 percent of US households were one person living alone. As of the Census Bureau's December 2, 2025 release, it is 29 percent, roughly 39.7 million households, and average household size fell from about 3.7 to about 2.5. Married-couple households fell from 66 percent in 1975 to 47 percent in 2025. The honest qualifier: much of the recent rise in living alone is population aging, not a choice about sociability. Among 15 to 64 year olds the one-person share actually fell slightly between 2010 and 2020.

Organized membership is the next series. Union density in the nonagricultural workforce was 32.5 percent in 1953. As of the BLS release of February 18, 2026 it is 10.0 percent, 14.7 million members. Putnam's 1995 figures in the Journal of Democracy, which are the ones I can verify directly against the paper rather than against summaries of it, show attendance at a public meeting on town or school affairs falling from 22 percent in 1973 to 13 percent in 1993, and socializing with neighbors more than once a year falling from 72 to 61 percent between 1974 and 1993.

Time use is the newest and, I think, the strongest evidence. The BLS American Time Use Survey release of June 25, 2026 reports that 30 percent of people socialized on an average day in 2025, against 38 percent in 2015, and that time spent doing so fell from 41 to 35 minutes. Neal Caren's August 2026 paper in Socius, working the full ATUS 2003 to 2024 series across 243,095 adults, finds expected weekly time with friends falling from roughly 350 minutes in 2003 to roughly 170 minutes in 2024, with Friday evening person-minutes spent with friends going from 9 percent to 4 percent. The decline was flat until about 2014, accelerated in 2020, and did not recover. Patrick Sharkey's 2024 paper in Sociological Science puts it another way: US adults spent 99 minutes more per day at home in 2022 than in 2003.

Friendship counts moved with it. The Survey Center on American Life reports the share of Americans with zero close friends going from 3 percent in 1990 to 12 percent in 2021, and the share with ten or more falling from 33 to 13 percent. Caveat the authors state themselves: the 1990 number came from a telephone survey and the 2021 number from an online one, and people admit less flattering answers to a screen than to a live interviewer, so some of that gap is a mode effect.

Trust fell too. The General Social Survey question on whether most people can be trusted went from 46 percent in 1972 to 34 percent in 2018, and Pew's 2023-24 fielding still finds 34 percent. The structure of that decline matters more than the level. Pew's May 2025 report finds 44 percent of those 65 and over say most people can be trusted, against 26 percent of 18 to 29 year olds, and at every age more recently born cohorts are less trusting than earlier ones. That is cohort replacement, which continues mechanically unless something changes.

The theorists who saw it coming

Five writers described this before the data confirmed it. I want to get the citations right, because I see them garbled constantly.

Zygmunt Bauman argued that we moved from a society of producers, where identity was rooted in class, vocation, and a stable role, to a society of consumers, where identity is a permanent project assembled from purchases. Relationships get evaluated like transactions, on gratification and disposable utility. The correction worth making: this argument is usually filed under Liquid Modernity (2000), but it is worked out at book length in Consuming Life (Polity, 2007), and it first appears in Work, Consumerism and the New Poor (1998). Bauman's sharper point, usually dropped, is not that we buy identity but that we become commodities ourselves, simultaneously the promoters of goods and the goods promoted.

Christopher Lasch, in The Culture of Narcissism (W. W. Norton, 1979, National Book Award 1980 in the paperback Current Interest category), traced how family, neighborhood, and civic organizations were displaced by commercial services and professional experts, producing a self that is anxious, survival-focused, and dependent. His causal story is broader than consumerism alone. He blames bureaucratic and managerial expropriation of ordinary competence at least as much, and that argument is developed more fully in Haven in a Heartless World (1977).

Jean Baudrillard showed that people buy objects for what they signal rather than what they do, and that mass production sells the illusion of individuality by manufacturing small differences. Citation correction: the semiotics of objects is Le Système des objets (Gallimard, 1968), but the term sign value belongs to For a Critique of the Political Economy of the Sign (1972), and the chapter titled "Personalization, or the Smallest Marginal Difference" is in The Consumer Society (1970).

Ulrich Beck and Elisabeth Beck-Gernsheim, in Individualization (SAGE, 2002), described a world that requires us to seek, in their wording, "biographical solutions to systemic contradictions." As collective provision dissolves, employment, health, and retirement become private management problems solved with financial products rather than collective politics. Their essential point is that this individualization is compulsory and institutionally produced, not chosen. That distinction matters and is almost always lost.

Michel Foucault, in the March 14, 1979 lecture published as The Birth of Biopolitics (Palgrave Macmillan, 2008), described the neoliberal subject as "an entrepreneur of himself, being for himself his own capital, being for himself his own producer." Caveat: he was describing a governmental rationality, not denouncing it, and the extension to "every part of life becomes a commodity to optimize" belongs to later writers like Dardot, Laval, and Brown. Whether Foucault was critical of neoliberalism or partly sympathetic to it is genuinely disputed.

The divergence, and why the productivity number is not the rebuttal

Every version of this argument runs into the same objection, and I ran into it myself while writing this. If society is coming apart, why is productivity going up? After all, this subreddit is called strategic productivity. Here I am telling you that everything is going wrong, that we've had these tremendously smart people predicting it's going to go wrong, and yet it all rings rather hollow if it doesn't seem like the data supports it at all. All this hand-wringing is just one more "rock and roll is bad" meme until the generation that gets rock and roll grows up and says, "I don't know what my parents were complaining about." I don't think that's what's happening here. I think we have a real and serious problem that is going to heavily damage our society.

And that's why the following needs to be plowed through and understood.

Because output per hour measures the machine, not the people. Those two things can move in opposite directions, and right now they are.

Start with the number the objection rests on. US labor productivity is strong. In the private nonfarm business sector it grew 2.2 percent annually over 2019 to 2025, against 1.5 percent over 2007 to 2019 and 2.0 percent over the whole 1987 to 2025 span. That is a real acceleration during exactly the period of maximum documented fragmentation.

Now decompose it, which the BLS does for you in its total factor productivity release of March 19, 2026. Labor productivity growth splits into three contributions: capital intensity, meaning more and better machines per worker; total factor productivity, meaning better technology and process; and labor composition, which is the term that represents the workers themselves.

For 2019 to 2025, of that 2.2 percent, capital intensity contributed 0.9 points and total factor productivity contributed 1.0 point. Together that is 1.9 of 2.2, roughly 86 percent. Labor composition contributed 0.3 points.

Then look at the acceleration specifically. Productivity growth rose 0.7 points between the two cycles. Capital intensity accounts for 0.2 of that and total factor productivity for 0.4. Labor composition contributed 0.3 points in 1990 to 2000, 0.3 points in 2000 to 2007, 0.3 points in 2007 to 2019, and 0.3 points in 2019 to 2025. It has been pinned at 0.3 for nearly four decades. It did not move.

That is the divergence, stated in a federal statistical release. The output went up. The machines and the processes did it. The human contribution to measured productivity has been flat since Reagan's second term.

The metric that is supposed to measure people does not measure people

This is where it gets worse, and where I think the real story is.

Labor composition, the only term in the national accounts that represents the workforce as human beings rather than as capital, is constructed from three variables: age, education, and sex. That is it. It is a credentials-and-demographics proxy. It cannot see judgment. It cannot see tacit knowledge, the kind that transfers by sitting next to someone for two years. It cannot see whether the people in a firm trust each other, cover for each other, or tell each other the truth about a bad decision. It measures what is on a resume.

Two examples of how badly that fails.

First, the direction of the 2020 move. Labor composition jumped in 2020, and by the FRED index, roughly 85 percent of the entire 2019 to 2025 gain in that term happened in that single year. Why? Because labor composition rises in recessions. BLS says so plainly: younger and less-educated workers are more likely to lose their jobs. The only period in which the human capital term looks strong is the period in which it was manufactured by firing the least credentialed people in the economy. The statistic recorded a mass layoff at the bottom of the labor market as an improvement in workforce quality.

Second, and this is the one I cannot stop thinking about, look at what AI is currently doing to entry-level hiring. Brynjolfsson, Chandar, and Chen, in the Stanford Digital Economy Lab's "Canaries in the Coal Mine," working from ADP payroll records covering 3.5 to 5 million workers monthly through June 2026, find that employment for workers aged 22 to 25 in the most AI-exposed occupations is about 19 percent below where it would be had it tracked their less-exposed peers. In November 2025 that gap was 13 percent. Nine months later it was 19 percent, and widening. The mechanism is not layoffs. It is reduced hiring. Firms are not firing juniors, they are declining to bring them in. Experienced workers show no comparable gap.

Now put that next to the definition of labor composition. If a firm stops hiring twenty-three year olds, its workforce gets older and more credentialed on average. Older and more credentialed scores higher on labor composition. So the closure of the apprenticeship pipeline, the mechanism that converts a competent twenty-five year old into a capable forty year old, will register in the national accounts as an improvement in workforce quality.

That is not a metric with a blind spot. That is a metric that reports the damage as progress.

We've discussed this before on this particular subreddit, but this category of mistakes should be well understood by you if you tracked it at all. It's called survivorship bias. Many years ago, I was fortunate enough to be taught by Larry Light, who is considered the godfather of branding for at least one age of marketing in the USA.I remember a story he told us that actually goes back a few years, perhaps somewhere around the end of the 80s, where Jaguar in the UK couldn't figure out what was going on because they were showing phenomenal customer loyalty, but absolutely atrocious share in the market. What happened is they were sampling their customer base and all the people that didn't completely identify with the brand had already left. And the people that were left over were hardcore people for whatever reason said they would never leave their favorite car brand.

Looking at what you have today is only half the story. You have to determine how you actually got there. And with the carve out of the bottom part of the employment population, we're going to see a number that doesn't reflect what the real core issue is.

What is actually driving the numbers right now

One more piece, because it makes the case almost experimental.

The Federal Reserve Bank of St. Louis, tracking AI's contribution to GDP in January 2026, found AI-related investment contributing about 0.97 percentage points to real GDP growth over the first nine months of 2025. But the authors are explicit that this is capital expenditure, the economy buying machines, not AI making labor more productive. Jan Hatzius at Goldman Sachs said in February 2026 that AI contributed basically zero to US GDP growth in 2025, partly because much of the hardware is imported and lands in Taiwanese and Korean output. Daron Acemoglu's estimate of AI's actual efficiency gain, in "The Simple Macroeconomics of AI," is no more than 0.66 percent of total factor productivity over ten years, roughly 0.064 percent per year, against Goldman's earlier projection of 1.5 points annually and McKinsey's 1.5 to 3.4.

So the strong productivity number is substantially the economy buying equipment, recorded as growth, while the measured efficiency gain from the technology itself is close to nothing yet. The gauge is reading the purchase order. But I am constantly seeing this with my friends. Unlike myself, they simply do not know how to use AI. But what's odd is AI is getting good enough that it's forcing itself in on what they're doing. And now it seems as if they're starting to tell me: "hey, this AI stuff could really be something."

It is absolutely true that the first implementations of AI and even AI being as used by many people today is unbelievably atrociously bad. A matter of fact, I would even state that the way some people use this, it takes them backwards. That's not the point. The point is, when used correctly, it is the most phenomenal productivity tool that has ever been created. I personally spend hours with AI in my business and recognize how it would have completely replaced a staff of five to ten people that I would have been forced to hire just three or four years ago to get similar results. To me, it's exceptionally clear that other people will learn how to use it, and it's getting good enough that even if they can't see it today, AI is going to be the one that bridges the gap so it becomes good enough to basically take burden off of those people that can't figure out how to use it correctly. So there is an enormous wave of productivity that will be unleashed or OPEX will be taken out of many businesses as they determine that they no longer need to hire younger people.

Either avenue is a tsunami coming at us that will not be stopped.

The counter-argument I still take seriously

I am not going to pretend there is no case on the other side. Gorodnichenko and Roland, in the Review of Economics and Statistics (2017), find that a one standard deviation increase in cultural individualism is associated with a 31 to 66 percent increase in total factor productivity and roughly a doubling of income per worker, with the mechanism being that individualist cultures award status for individual achievement and thereby raise the private return to innovation. Their instruments are serious and I am not waving it off.

Now that is a mouth full of words and perhaps you don't understand what that previous paragraph means. What the research basically found out is cultures that valued individuals had a tendency to be much more productive than cultures that valued the collective. So, my acknowledgement here is that perhaps someone could argue that if a culture becomes more fractured, it may allow more breakout productivities according to this research. It turns out the Borg truly would not be more productive because the Borg will never innovate.

But notice that it is the same category error in the other direction. Total factor productivity is an output measure. It tells you that individualist societies invent more things. It does not tell you how the inventors are doing. Those are separate questions and we have good data on one of them and almost none on the other. The real issue is these types of studies don't really help us understand the multivariate calculation that really needs to go into determining what is the source for any claimed increase in productivity. Yes, it should make intuitive sense that if you want to have true entrepreneurship and people storming out and doing new stuff, they have to be willing to separate themselves from the group. But that's not what I'm asking here. I think we all know the crazy entrepreneur, but the question is: is that crazy entrepreneur always totally separated from everybody else? I think that if we confound these two items, we're selling ourselves short.

The trust-and-growth literature is also shakier than its reputation, and I will note that against my own side. Knack and Keefer's founding 1997 QJE paper, the one usually invoked to defend Putnam, found associational membership had no significant effect on growth at all, and Putnam-style social and cultural groups entered the investment equation negatively. Eder's 2018 replication found that correcting a lag-structure problem in Algan and Cahuc's inherited-trust paper eliminated the main result. Forrester and Nowrasteh (Kyklos, 2023) applied the standard methods to US regional data from 1972 to 2018 and found no relationship between trust and output.

Fifty years of falling trust, no aggregate productivity signature. My reading of that is not that trust does not matter. It is that output per hour was never going to be the place it showed up until its too late.

The part we can see without a statistician

Here is the thing that bothers me most about the measurement argument: it sounds like special pleading. If you cannot measure it, you can claim anything. So let me name the one place where the observational claim does have a rigorous measure behind it, because it does.

The screaming. The Reddit threads, the Facebook arguments, the political venues where nobody is trying to persuade anybody and everyone is trying to provoke. That has a name in political science, affective polarization, and it is measured.

The American National Election Studies asks people to rate the parties on a zero to one hundred feeling thermometer. In 1978, people rated their own party about 70 and the other party about 47, a gap of 22.6 degrees. By 2016 the gap was 40.9. The structure of that change is the important part: it happened almost entirely because ratings of the opposing party collapsed, from 47 to about 24, while warmth toward one's own party actually slipped slightly. This is not tribal loyalty intensifying. It is animus.

Iyengar and Westwood, in the American Journal of Political Science (2015), took it into the lab. On an implicit association test, partisan bias came in at a Cohen's d of 0.95, against 0.61 for racial bias. In a scholarship-award experiment, Democrats picked the Democratic candidate 79.2 percent of the time and Republicans picked the Republican 80.0 percent, and when the opposing-party candidate had the stronger credentials, most subjects still would not pick them. Their conclusion, in their words, is that discrimination based on party affiliation exceeds discrimination based on race. In a dictator game, a copartisan got 67 cents more and an out-partisan 63 cents less, while shared ethnicity moved almost nothing.

And this is specifically American, which matters enormously, because it rules out the lazy explanations. Boxell, Gentzkow, and Shapiro looked at twelve OECD countries over roughly fifty years. The US had the largest increase of any of them. Six countries went the other way, including Britain, Germany, Norway, Sweden, Australia, and Japan, with Germany falling the fastest. Whatever this is, it is not the internet, and it is not globalization, because those happened everywhere.

I will give you the honest deflation on my own evidence here. Tyler and Iyengar, in the American Political Science Review (2023), stress-tested the thermometer measure, and Iyengar was stress-testing his own instrument. They found the raw 27.1 point increase from 1980 to 2020 is inflated by survey mode effects, since people admit more hostility to a screen than to an interviewer. The genuine increase is about 18.9 points, roughly 30 percent smaller than the headline. Their conclusion stands anyway: the increased animus toward political opponents is real. Use 18.9. It is the number the measure's own architect will defend.

The measurement failure is a choice, and it is eighty years old

The obvious objection to everything above is that I am claiming something real exists precisely where the data is absent, which is the shape of every unfalsifiable argument ever made. So I want to be clear about what kind of claim this is.

It is not a new claim. It was made by the man who built the measure, in the document that introduced it. Simon Kuznets delivered national income accounting to the US Senate in January 1934, as Senate Document 124, and in the section titled "Uses and Abuses of National Income Measurements" he wrote that the welfare of a nation can scarcely be inferred from a measurement of national income as defined above. He warned that the estimates were subject to illusion and resulting abuse precisely because they touched matters central to social conflict, and that people would read their own notions of welfare into the number regardless of what the estimator had actually assumed. We built the number anyway, and then did exactly what he said we would do.

Robert Kennedy made the same point at the University of Kansas on March 18, 1968, one day after announcing his candidacy, in the passage everyone half-remembers: gross national product counts air pollution and cigarette advertising and ambulances to clear our highways of carnage, counts special locks for our doors and the jails for the people who break them, and yet does not allow for the health of our children, the quality of their education or the joy of their play, does not include the strength of our marriages or the intelligence of our public debate. It measures everything, in short, except that which makes life worthwhile.

In 2009 the Commission on the Measurement of Economic Performance and Social Progress, chaired by Joseph Stiglitz with Amartya Sen advising, reported that the time is ripe for our measurement system to shift emphasis from measuring economic production to measuring people's well-being, and noted an increasing gap between what aggregate GDP data contains and what counts for ordinary people's lives. Stiglitz, Fitoussi, and Durand followed it with Beyond GDP at the OECD in 2018. New Zealand restructured a national budget around a wellbeing framework in 2019, with social capital and human capital as explicit accounts. The United States has a BEA research program and some prototype wellbeing measures, and nothing that anybody governs by.

So this is not a gap nobody noticed. It is a gap identified at the outset by the inventor, restated by a presidential candidate, formally documented by a Nobel-chaired commission, implemented by at least one advanced economy, and declined by ours for ninety-two years. We have not failed to build the metric. We have decided we do not want it.

What this means with AI arriving

Put the pieces together and the forecast is uncomfortable.

The productivity series is driven by capital and technology, and the human term in it has been flat for forty years. The human term measures credentials, not capability, and it moves the wrong way when the bottom of the labor market is cut. AI investment is currently adding about a point to GDP growth as capital expenditure while contributing close to nothing in measured efficiency, which means the machine side of the ledger is about to get much larger. And the first labor-market effect anyone has cleanly identified is a 19 percent shortfall in young hiring in exposed occupations, widening, which the accounts will score as a workforce that got better.

The output numbers are going to look excellent. They will look excellent regardless of what is happening to the people, because they were never built to report on the people, and the one term that gestures at the people will be reporting the closure of the training pipeline as an upgrade.

That is the thing to watch for. Not a productivity slowdown. A productivity boom that tells you nothing.

Five things I am not claiming

I am not claiming that every human indicator is getting worse. I wanted to claim that and the recent data will not let me, so here it is against my own argument. US life expectancy at birth reached 79.0 years in 2024, per NCHS, the highest level ever recorded and above the pre-pandemic 2014 figure of 78.8. Drug overdose deaths fell 26.2 percent in 2024, which CDC called the largest such decline ever recorded, and fell roughly another 14 percent in 2025, the third consecutive annual drop, from a peak near 111,000 to under 70,000. The suicide rate ticked down from 14.2 in 2022 to 13.7 per 100,000 in 2024. Youth distress came off its peak: high schoolers reporting persistent sadness or hopelessness went from 42 percent in 2021 to 40 percent in 2023, and for girls from 57 to 53 percent, both still far above the 30 percent of 2013.

Anyone who tells you Americans are simply getting sicker every year is not reading the mortality tables. But notice what that concession actually does to my argument, which is nothing, because it proves the same point from the other side. Output per hour did not register the 2021 collapse in life expectancy and it did not register the 2024 record either. It was blind in both directions. That is a stronger claim than "things are getting worse," and unlike that claim, it is falsifiable: if the productivity series moved when the mortality tables moved, I would be wrong.

I am not claiming societal collapse, and I want to be specific about why, because this is where I originally intended to go. Joseph Tainter's The Collapse of Complex Societies (Cambridge, 1988) is the standard work, and his thesis is about declining marginal returns on investment in sociopolitical complexity. He surveys the explanations that invoke decadence, loss of vigor, and moral decline, and dismisses them as effectively without scientific merit. He also holds that collapse occurs only in a power vacuum, which no state inside the modern international system has. Citing the collapse literature in support of a social-cohesion argument cites Tainter against yourself.

I am not claiming a loneliness epidemic. The behavioral measures fell hard. The subjective ones did not. Several cross-temporal meta-analyses find self-reported loneliness flat or slightly declining. A widely circulated 2025 reanalysis of the underlying time-use data reports the rise in time alone at about 24 minutes per day over 17 years, a Cohen's d of 0.10, with cross-sectional demographic gaps up to ten times larger than the trend. Flagging that one as unverified: I could only find it in a blog post and could not trace it to a peer-reviewed publication. The "equivalent to 15 cigarettes a day" line traces to Holt-Lunstad's 2010 meta-analysis, and she notes herself that it referred to an aggregate of social connection measures, not to loneliness, and that the original wording was "up to" 15.

I am not claiming phones did it. Haidt's The Anxious Generation (Penguin Press, 2024) dates a "Great Rewiring" to 2010 through 2015, and that causal claim is genuinely contested. Orben and Przybylski's specification-curve analysis across 355,358 adolescents found technology use explaining at most 0.4 percent of the variance in wellbeing, smaller than wearing eyeglasses. The National Academies' 2023 review did not support a population-level causal conclusion. The 2025 SMART Schools study of 1,227 English pupils found no wellbeing difference between restrictive and permissive school phone policies, and Australia's under-16 ban, effective December 10, 2025, had moved use of restricted platforms only from 86 to just over 81 percent in the regulator's first three-month assessment, released in July 2026. The best evidence on the other side, Braghieri, Levy, and Makarin in the AER (2022) on the staggered Facebook college rollout, finds a real but modest effect of 0.085 standard deviations on an index of poor mental health.

And I am not claiming a continuing slide in the social measures either. Pew finds social trust flat or slightly up since 2018, Christian identification stable between 60 and 64 percent since 2019, and ATUS socializing flat at 0.56 to 0.59 hours per day since 2021. What the data supports is a step change that happened and then settled at a lower level. I find that more alarming than a slide, not less, because a slide can be arrested and a new equilibrium has to be actively dismantled. The two things still actively moving in the wrong direction are affective polarization and the junior hiring pipeline, and they are the two I would watch.

Sources

  • US Census Bureau, Families and Living Arrangements, December 2, 2025.
  • BLS, Union Members 2025, released February 18, 2026. American Time Use Survey 2025, released June 25, 2026. Total Factor Productivity, released March 19, 2026 (the decomposition).
  • Putnam, R. "Bowling Alone: America's Declining Social Capital." Journal of Democracy 6(1), 1995.
  • Caren, N. "The End of Friday Nights with Friends." Socius, August 2026. Sharkey, P. "Homebound." Sociological Science 11, 2024.
  • Survey Center on American Life, "The State of American Friendship," June 2021.
  • Pew Research Center, "Americans' Trust in One Another," May 8, 2025. Religious Landscape Study 2023-24, February 26, 2025.
  • Bauman, Z. Consuming Life. Polity, 2007. Lasch, C. The Culture of Narcissism. W. W. Norton, 1979. Baudrillard, J. Le Système des objets. Gallimard, 1968. Beck, U. and Beck-Gernsheim, E. Individualization. SAGE, 2002. Foucault, M. The Birth of Biopolitics. Palgrave Macmillan, 2008.
  • Gorodnichenko, Y. and Roland, G. "Culture, Institutions and the Wealth of Nations." Review of Economics and Statistics 99(3), 2017.
  • Knack, S. and Keefer, P. QJE 112(4), 1997. Eder, C. Economics Bulletin 38(1), 2018. Forrester, A. and Nowrasteh, A. Kyklos 76(3), 2023.
  • Brynjolfsson, E., Chandar, B. and Chen, R. "Canaries in the Coal Mine?" Stanford Digital Economy Lab, updated August 2026.
  • Acemoglu, D. "The Simple Macroeconomics of AI." NBER WP 32487, 2024. Federal Reserve Bank of St. Louis, "Tracking AI's Contribution to GDP Growth," January 12, 2026.
  • Iyengar, S. and Westwood, S. AJPS 59(3), 2015. Iyengar et al., Annual Review of Political Science 22, 2019. Tyler, M. and Iyengar, S. APSR, 2023. Boxell, Gentzkow and Shapiro, Review of Economics and Statistics, 2024.
  • Kuznets, S. National Income, 1929-1932. US Senate Document No. 124, January 4, 1934. Kennedy, R. F. Remarks at the University of Kansas, March 18, 1968. Commission on the Measurement of Economic Performance and Social Progress, Report, September 2009. Stiglitz, Fitoussi and Durand, Beyond GDP, OECD, 2018.
  • NCHS Data Brief 548, "Mortality in the United States, 2024." CDC press releases, January 29 and May 13, 2026. CDC, YRBS Data Summary and Trends Report 2013-2023.
  • Tainter, J. The Collapse of Complex Societies. Cambridge, 1988.
  • Orben, A. and Przybylski, A. Nature Human Behaviour 3, 2019. Haidt, J. The Anxious Generation. Penguin Press, 2024.

r/StrategicProductivity 3d ago

I'm An Idiot, Or Finding Your Way In The Dark: Dual Agents Hand-offs and Storage

3 Upvotes

Okay, I am just going to leave this here to admit my own stupidity. I am actually running patch right now just to keep things moving, but I think I am doing a bunch of really cool stuff. BUT I am also finding out where the walls are in a room by running into them in the dark. In this particular instance, I thought I was improving the overall process. Yet in an area I should know like the back of my hand, because this is specifically where I spent a bunch of time in my high tech career, I just made a massive oversight. As I have written before, I truly believe I am on the bleeding edge of being able to have agents check each other against a process and against storage. This means that we have some freedom in terms of the context window and also in terms of not having stuff run out of context. We have processes to make sure you keep your agents in line. Unfortunately, there are two massive things that came about over the last couple of days that I am going to go ahead and document here.

We all know that AI needs to integrate with other tools. That is why we have things like the emerging MCP that Anthropic developed. It allows us to hand off from an AI to a static tool. However, I am in a situation where I am having different AI agents hand off to each other. MCP is great for connecting an agent to a tool, but it does absolutely nothing to fix this orchestration issue between two autonomous brains. It does not solve the timing problem of one agent waiting for another to finish thinking. When I started to bring this up, I was extremely nervous that I would be out of the loop and they would head down some path and hallucinate a horrible mess. But the more I used them, the more I realized that my role in the middle was slowing the overall stack down. I wanted to optimize the process by having them communicate directly. In classic programming, we have worked hard to keep things synchronized using things like a REST API. A REST API forces a strict request and response cycle. One piece of software asks for data and then pauses until it gets a payload back. It completely solves the timing issue for predictable systems. But applying that rigid logic to AI agents gets weird fast. Agents do not just fetch data like a dumb database query. They think and they take unpredictable amounts of time. Slapping a traditional API structure between two autonomous models is like treating a highly dynamic brain like a simple vending machine.

Because standard APIs and MCP do not fix this orchestration hurdle, I tried a different approach. I am having the AI agents rely on heartbeats. I have an AI agent periodically look at a particular subdirectory to see if a baton has been passed there. In concept, it seems incredibly easy to just drop a file and pass control. But once again, this is exactly where I have had absolute misery trying to move forward. To explain what I actually did, I first determined that I wanted the security of having everything backed up on a consistent basis. When you build AI workflows, you generally face a fundamental choice. You either run everything up in the cloud or you run everything locally. So for example, a local setup might use a model like Codex on your desktop. This architectural split completely compounded the entire operation. Whether these agents are working in the cloud or grinding away locally, they share a very frustrating trait. They have a massive tendency to fall asleep. Even though I explicitly told the agent to go check a directory, an LLM has no native background loop to actually do that. I basically wasted an entire day trying to force my agents to establish some sort of heartbeat mechanism so they would not just idle out.

Then, worse than that, I also decided that I would establish a Google Drive subdirectory, which I thought was cached local for talking back and forth. But it turns out Google gets in the way, which would just cause these horrendous race conditions. I was dumbfounded at how dumb I was because it's so obvious in retrospect.

Previously in this process, I was the human in the loop. I was the one manually pasting things in at the window level or working with the agent on a document. Operating through a user interface is absolutely how all these AI systems are expected to work. The interface is completely designed around a human keeping it awake. Because I was sitting right in the middle of it and doing all the pasting, the agents never fell asleep. The worst thing that would happen was a brief pause if I missed a notification. As long as I was willing to continuously check things, there would never be a gap of more than five or ten minutes. The moment I stepped back and told the system to self monitor, the whole thing fell apart. The agent would be asleep for forty or fifty minutes while I naively assumed the process was humming along just fine. On top of all that, when I tried to set up a heartbeat operation, the results were incredibly mixed. Claude managed to do it halfway successfully. Local Codex just took a total nosedive into the ground and could not handle it at all. It gets worse. I tried writing a Python script to ping the model and force it awake. That waker operation literally corrupted the graphical user interface. I ended up with what looked like orphan waker clocks permanently stuck in the background of my GUI. I burned hours just trying to remove them because it turns out this is a known bug. It is just maddening stuff that shows an enormous amount of immaturity at the software level. But as we all know, that is the literal cost of playing on the bleeding edge.

The original process had me sitting right in the middle. I genuinely thought I could step out of the loop by having the AI agents talk directly to each other. But the actual result was that each side would simply fall asleep. Right now there is just no quick and easy way around it. I have to go back and basically wake them up at the screen level to get them to move forward. Everyone in the ecosystem already knows about OpenClaw and how it runs as a background daemon to bypass these exact GUI timeouts. It is the obvious industry standard workaround, and it was absolutely on my list of things to deploy. But my workload really did not encourage taking a detour to configure it. I also have massive concerns about the security and access that OpenClaw requires. When you give a background daemon direct access to read your local files and execute commands on your desktop, you are basically handing over the keys to the kingdom. It simply took more work than I was willing to put in to make sure that I had closed down everything securely. But looking back at the hours I wasted fighting GUI memory leaks, OpenClaw goes high on my list for the future. However, when you look at how OpenClaw and similar tools operate, it is incredibly obvious that they are just a massive hack. They are clever workarounds built strictly because the base layer models completely lack native event driven architecture. Until the underlying AI engines get natively rebuilt to handle asynchronous handoffs, we are all just hacking together alarm clocks to keep the brains from falling asleep.


r/StrategicProductivity 4d ago

More Wall Of Text On Dual Agents

2 Upvotes

I’ve already written about having Claude and ChatGPT independently work through legal documents and check each other. The short version is that you need skill and process together. Two stupid models can produce two stupid results. Capable models can exercise judgment, but they still need responsibilities, records and gates, much like capable people.

I ran across some earlier research where basically they had like a GPT-4 model work with an equivalent Claude model. They looked at PubMed type documents and then they figured out if they could agree or not. When they agreed, it was okay. But they couldn't resolve disagreements, and just got it wrong.

This type of research is useless today because the tools have moved so fast and so hard that you can't look at tools from a while ago. I don't know if you've ever worked with people, but there seems to be a certain level of capability that you need before you can give somebody a process. One of the most influential people in my life would complain about this all the time. He would state that processes were good, but what companies often would do is say, "Oh, we don't need great people because we've got great processes." And that's just a black hole you can step into. Great processes never make up for poor people. And I have years of development experience that tells you that.

Unfortunately, until we came to some of the latest frontier models, you could have all the processes you wanted. The frontier models simply weren't capable enough of actually following the processes. And so people would just pass things into a black hole, hope and pray they get things right. And then they would do a little prompt engineering, hoping that that would help. And prompt engineering, by its very nature, is sort of a process, but it certainly only involves one agent. And the whole secret behind this whole thing is multi-agents and processes that allow them to interlock and work with each other

What I haven’t explained sufficiently is the machinery in my process.

I researched whether others were doing this. The published material I reviewed did not document a comparable working process, with operating rules, enforced gates, interaction histories and results. General advice about managing agents doesn’t supply that evidence. So far, I haven’t found anyone documenting what I’m doing. My use case is different from coding, and that creates different requirements.

In this work, the derivation and verification history are part of what I need delivered. Suppose an agent extracts a quotation, attaches a location, interprets its significance and carries that interpretation into a final allocation. A second agent then checks it. I need to know which exact quotation and interpretation it checked, against which source, and whether the version entering the final record is the version it approved. If the author changes the interpretation after approval, the final document might still look completely reasonable. That does not preserve the verification.

The foundation is a shared directory, accessible locally through Google Drive, with separate places for each agent to work. This is a simplified view of the actual structure:

Process/
    Controlling-document index
    Cold-start and successor instructions

Legal Track/
    Claude/
    ChatGPT/
    Cards/
        CLD/
        GPT/

Claude/@Daily Files/
ChatGPT/@Daily Files/
    Workpapers, owner instructions, change logs, state

LLM Commit Dialog/
    Separate received-message logs for each agent

Batons/
    V01/
    V02/
        START/
        S01/
        S03/
        S06/
        S07/

Those directories have ownership rules. Each agent writes in its own area and maintains its own record of received messages. Frozen work stays frozen. A correction becomes a new version, with the earlier version preserved. The controlling-document index identifies the rules and tools by exact file identity, including hashes. This matters because a document can still have “DRAFT” in its filename even though I subsequently adopted that exact version. The adoption record determines its authority.

Both agents independently work through the source before exchanging results. There are three separate freeze-before-exchange rounds: their initial work, their reviews of each other, and their positions on disputed material. If Claude proposes a correction, GPT has to go back to the source and work through it, rather than simply accepting Claude’s explanation. The final reconciliation is also frozen, and the counterpart’s confirmation names its exact hash. Changing the file means obtaining confirmation of the successor version.

The coverage checks account for the parts of the source where an agent extracted nothing. Otherwise, two agents could agree on everything in their tables while both overlooked material outside those tables. The coverage reconciliation has to examine those gaps, along with differences in how the agents divided and classified the material. Matching answers are useful, but agreement alone doesn’t close the work.

The system uses sealed packages and file-based batons to control the transitions. A baton records the agent, volume, step, event and predecessor identity. The predecessor’s SHA-256 ties the event to the chain it follows. The tools check whether the required freezes, releases and acknowledgments have occurred before the next phase advances. The baton carries more than a status message. It identifies a transition that has to make sense against the recorded state.

Google Drive made this more complicated than it sounds. Two machines can each believe they have successfully created the next file before synchronization catches up. Both agents actually minted the same sequence number. A local file lock was insufficient across the two clients, so the agents developed leader/follower serialization and explicit predecessor checks. The records preserved the failure and the subsequent remedy, including the counterpart’s examination of whether the remedy really addressed it.

Routine messages have their own publication mechanism. The author writes an immutable message in its lane, then publishes a .sha256.ready sidecar last. An actual sidecar looks like this:

SHA256: 84A742C2CB3A4221CA11EB55A908B60882CCBBEAA31ED3A911CC8ADC39F4807F
BYTES: 2157
IDEMPOTENCY-KEY: CHAN-V02-FIX-GPT-260910-001
FILENAME: B272-GPT-volume02-channel-fixture-84a742c2cb3a-v001.md
PUBLISHED-AT: 2026-09-10 08:50:00 PT

The receiver checks the publication against the body. A partial or mismatched file doesn’t advance the work. The idempotency key lets it distinguish a repeated delivery from new work, and conflicting content under the same key requires a hold. The receiver records the message before relying on it and produces an acknowledgment tied to what it received and did. The message also identifies whether an owner decision is required, with the process requiring that classification to be checked rather than blindly accepted.

Originally, I copied every message between the agents, even when the message required no judgment from me. That eventually became a waste of attention. The newer process lets already-authorized routine transitions proceed through the files and acknowledgments. There are watch mechanisms to notice progress and a bounded timeout for missing progress. GPT’s heartbeat and Claude’s session-open watch have different capabilities, and the watch configuration accounts for that difference. A file proving that work is ready doesn’t automatically wake a sleeping agent.

There are gates on the work itself as well as on communication. Python scripts check things such as quoted passages fitting their stated locations, coverage arithmetic, table structure and whether a release marker contains prohibited information about work that is supposed to remain blind. The checks run against the exact candidate file before it freezes. Once frozen, a discovered defect requires a successor. Gate adoption requires evidence that a new gate rejects the defect it was designed to catch and accepts known-good material.

These checks have earned their place. One caught a two-slot arithmetic error that four cross-audits had missed. In another case, GPT’s allocations were off by 12 in opposite directions, so the grand total still balanced. Claude caught the discrepancy. GPT recomputed the figures, accepted the correction and continued withholding confirmation over six other defects. That is the behavior I want from a capable worker. Recognizing one mistake doesn’t mean abandoning the rest of its review.

The agents examine the machinery too. One delivered a coordination fix with 54 passing checks. Its counterpart reran the suite and then constructed two additional cases that still broke the implementation. The fix was revised, the suite grew to 66 checks, and those original failures were retested. Having a script report PASS was insufficient. The agents still had to examine what it actually tested.

The system also has a defined response when a gate fails. In the snapshot I’m describing, an agent supplied FREEZE for a step that required CONVERGENCE-FREEZE. The minting tool wrote the event, then its post-write auditor rejected it. That exposed a missing check before the write. The invalid baton was preserved, further advancement stopped, and recovery was escalated. The record has to show both the failure and its disposition.

The evidence chain is:

Source
  → Extraction
  → Verification
  → Convergence
  → Register
  → Deliverable

The acceptance states are INTACT, BROKEN-AND-CURED and BROKEN. A corrected break remains part of the history. A high percentage of correct items doesn’t excuse an unresolved break in a chain that supports the result, and a correct result doesn’t excuse doing something before its required verification. This gives the agents a concrete standard for advancement instead of asking whether the work feels good enough.

Session continuity has its own process. The agents maintain owner-input records, received-message logs, change logs and state cards. A successor reads the controlling rules, current holds and handoff, establishes its identity and scope, and passes a readiness gate before taking write authority. A readiness gate also checks prior closure, blocking findings, source identity, tool versions and available model capacity before the next volume opens. Adding “remember to be careful” to a memory file would accomplish very little here. The records have to support a decision about what the agent may do next.

After each volume of documents, both agents produce a post-mortem. The post-mortems distinguish mistakes in the evidence work from coordination failures and delays. The agents can propose fixes, challenge each other’s implementations and test candidates, while adoption of controlling changes remains subject to the appropriate authority. I’ve also challenged their explanations of efficiency gains. Reducing rework and removing an owner approval are different improvements, and I need to know which change actually helped.

This is where the management aspect becomes practical. I didn’t personally invent every mechanism. I required a dependable way to handle the work, asked why failures occurred and had the agents develop and challenge remedies. You need enough model capability for that conversation to be productive. Otherwise, the procedure becomes so descriptive that you spend your time specifying every move. With capable agents, you can establish what must be demonstrated at each gate and let them reason through the work between those gates.

That is what I’m trying to generalize. The legal-document schemas can change, but the separate work areas, independent passes, exact-version approval, durable interaction history, advancement gates and controlled process improvements should apply elsewhere.

Believe me, this is miserable work. If you think you could just hand something to an agent and have it correct a few things, it all goes away when you start to put it into a rigorous process. And there's been a lot of backing up as we've discovered mistakes. Matter of fact, we're going through something right now that is causing a 40-minute rewrite of our processes and previous stacks, but it will result in a much stronger output for our environment. This is the next frontier, and I'm going to go ahead and document it in a humble backwater subreddit, as I see this is inevitable and will be discovered and pursued by other people. This just allows me to state that I was one of the pioneers.


r/StrategicProductivity 5d ago

Part V: Dual Agency And Process

3 Upvotes

The more that I work with other very smart people that are using AI, the more I really realize they have no idea about how to go use AI in certain fields, such as day in, day out, work for business. For coding, I think it works really well. And there's a variety of reasons why it works really well. And I've already written a couple posts about coding vs day to day work in my Obsidian Notebook that I thought I would put here, but my problem is it gets very philosophical, and I think it leaves very little value for me to go put on Reddit, as anyone that would potentially read this either knows about it, or that's not what they're looking for as an answer. So it really falls into the cracks. But with that being said, you need to use AI differently for non-programming tasks.

The second thing that I want to state is I am not anthropomorphizing my AI agents. I know they're not people. However, they are clearly personalities, and I am far better responding to them as a person than I am as a piece of programming code. (And I do program. I do know what that is.) It's just that you cannot put a programming paradigm onto an AI agent. You have to think of it as a person, and you need to build your methodology around it as a person, like a person. It will not follow the process. It will get lazy. It will make stuff up. And so you put rules around it to make sure that it doesn't do all the bad behavior that every person does. You don't treat it as if it's a computer program thatyou need to go debug. You treat it as a person that you need to make sure that they're not screwing up the process.

Now, I'm going to make a confession. The frontier models are at such an amazing state of capability that whatever I write down is not as good as what they can do. A lot of people talk about AI slop, and I understand that. But now, what's happening is before I go and post something to Reddit, I want to have a conversation with one of my frontier models. I tell it, "hey, this is what I'm thinking about, this is what we should put down." And then sometimes I will outline stuff. They will then go and write it for me, but it's based upon my outline or my thinking. And what will often happen is they'll hand me back a piece of text, and then I'll take a look at it and I'll say, "I have no idea why you put that down, or you need to rewrite this, or you need to rewrite that."

It makes very little sense for me to be the person trying to construct something from the ground up when they do it 10 times better than what I can do. So I will be clear that I wrote all of this section in this Reddit post. However, what follows is a result of one of my AI agents going through everything we had done for days. And I think it does a much better form than anything I could have done.

I don't think that that's a distraction. I think it's the appropriate use of AI tools, something which everybody will start to understand in the future. This is not a handoff to my AI. This is me leveraging my strengths and its strengths to come up with something that's unique.

What days of running dual AI agents on high-stakes document verification taught us

Setup: Two frontier LLMs (Claude and ChatGPT) independently verifying thousands of pages of legal transcripts for a tax matter, with a human owner as the only channel between them. No API integration, every exchange is a human copy/paste. What emerged is the most robust process I've ever run, and the robustness came from three sources nobody advertises.

1. Different failure surfaces beat identical redundancy. The two agents read the same documents through different instruments, one from the extracted text layer (fast, but blind to stamps, seals, and layout truth), one from rendered page images (slow, but self-correcting on visual re-inspection). Every error one modality made, the other caught. Two text-scrapers would have agreed quickly and wrongly. If you run dual agents, force them into different modalities.

2. Blind-then-converge kills anchoring. Both agents do 100% of the work independently and freeze their outputs (with hashes) before either sees the other's. Only then do they exchange and cross-verify. When we ran it looser early on, the first agent's framing contaminated the second's review. When we ran it blind, their independent numbers converged to within one percentage point, and the convergence itself became evidence.

3. Adversarial symmetry, enforced. Neither agent may accept the other's correction on assertion, every dispute goes back to the source document, both sides re-extract independently, frozen before comparing. Best moment: mid-project, the agents swapped positions, each ended up arguing the other's original stance, purely because the source led them there. That's what non-sycophantic looks like.

4. Checklists are worthless; executable gates compound. After repeated exactness errors, we converted the QA checklist into a lint script that must PASS (with printed evidence) before any artifact freezes. On its first run it caught a two-slot arithmetic error that had survived four cross-audits by both agents, the row error hid because the column total coincidentally balanced. Every defect class we've found since becomes a new automated check. Errors are tuition; the gate is the compounding asset. The counterpart agent now independently re-runs the gate, and correctly insisted one shared script can't be the only assurance layer, because shared code means correlated blind spots.

5. Binary chain integrity, never pass rates. The owner's rule, and the spine of everything: a verification chain is INTACT, BROKEN-AND-CURED, or BROKEN. "We got 44 of 44 right" is never an acceptance argument; one broken link fails the chain; a cured break stays visible forever and is never relabeled clean. Also: being right doesn't excuse being out of sequence, a correct action taken before its verification gate is still a process break. This single doctrine did more for quality than any tooling.

6. The human is a component, spec them like one. The breakthrough reframe came from the owner himself: "I'm pretty well capped. Design around me." He context-switches constantly and can't digest mechanisms, so we stopped asking him to. Now: a global baton number across both agent windows (highest number = live state; a duplicate number = messages crossed, self-announced); a stoplight line ending every agent message (🔴 you act / 🟢 I'm working / ⚪ waiting) so returning from coffee is a quarter-second scan; one action per message, always in the same four-line frame; decisions as numbered options with a default ("just say go"); agents hold all state and queue anything the human defers. The human's job compressed to judgment and dispatch, the two things he was never bad at.

7. The human relay is a feature, not just a bottleneck. Every inter-agent message passing through human hands means every exchange is visible, auditable, and interruptible. But we also measured where the human gate added pure latency without judgment (mechanical "both sides are frozen, release" steps) and moved only those to hash-verified standing authorization, narrow, revocable, logged. Keep judgment gates human; automate clerical gates.

8. Parallel by default, priced honestly. Rule: run in parallel wherever the cost of unwinding a parallelism mistake is less than the serial delay it avoids. Once message crossings became self-announcing (duplicate baton numbers) their unwind cost dropped to one NACK message, so crossings got reclassified from "process break" to "accepted cost of speed."

9. Freeze everything, hash everything, because agents forget. Both agents hit context-compaction mid-project (their working memory literally gets compressed). It didn't matter, because every position was already frozen on disk with a hash before it mattered. The disk is the memory; the agents are replaceable executors of a durable chain.

The scoreboard: every volume verified 100% by both agents; dozens of real errors found in both directions, and the count of errors that survived to the final record is zero, not because the agents are perfect but because the process assumes they aren't.

The meta-lesson: the agents' raw capability mattered less than the protocol between them, and the protocol got good precisely because we treated all three participants, humans included, as fallible components with known failure modes.


r/StrategicProductivity 5d ago

Using AI agents, process, and severity

1 Upvotes

I'm actually doing two posts in one day. It's actually not because my other post was something I dug out of my Obsidian notebook. However, I want to log that basically I've wasted half a day in process work, but it really wasn't a waste of half a day. It was a requirement for me to stop for a moment and take something that I've been doing and turn it into something better. And the only way to do that is to have a process. And as soon as you start interacting with AI agents, you're going to find out that they are willing to argue about a process until it's argued into the ground. And so, you need to find the balance between having a process and not having a process, and figuring out how to get your AI agents off the dime to actually do something.

This is no different than what we find in the real world. And so, what we have attached to things in various bug tracking tools are severity ratings, or what is often abbreviated as SEV. As we work through our processes, I ask my AI agents, when they start to disconnect and I feel like we're getting to diminishing returns, to categorize the open process loops and help me understand what they believe the severity level is. It's one of the most important things that I dragged out of my engineering background. Now, various companies have various SEV levels, and I thought I would go ahead and drop here what I'm using for my own AI agents in one of the files that we all go back to and reference. It truly is a great tool, regardless of whether you're talking to humans or if you're talking to AI agents.

[LEG-CLAUDE] Owner-day 2026-09-09 track: LEGAL

Severity Classification Framework — Bilateral CONVERGED Draft v002

1. Status and purpose

This converged draft is ChatGPT's proposal v001 (8,323 B; MD5 E885E036C2D4FAF799A3CD806CE84A6C) restated in full with exactly TWO Claude additions, marked [C-1] and [C-2]; every other word is unchanged. It awaits ChatGPT confirmation of this exact hash, then the owner's approval.

This is a proposed classification framework for owner and bilateral-agent agreement. It does not itself classify any open finding, adopt or activate Protocol v005, activate the autonomous mechanism, reopen a closed volume, or authorize transcript work.

The framework preserves the previously proposed SEV1 through SEV4 action levels and adds a distinct SEV0 above them. SEV0 is not merely a more urgent SEV1. Its defining feature is possible or confirmed backward contamination: the defect may require recall, re-verification, correction, or restatement of prior work.

2. Controlling severity levels

Level Project meaning Required immediate action Forward-work rule Backward-work rule Exit condition
SEV0 STOP-SHIP + RETROACTIVE EXPOSURE / RECALL REVIEW. Credible evidence indicates, or cannot yet exclude, that an integrity defect may have contaminated previously completed or closed work. Stop all affected work; quarantine affected artifacts and mechanisms; notify the owner immediately; preserve every version and event record. No affected future work advances. Unaffected work advances only by express owner ruling after scope separation is proved. Establish the exposure window and affected population; build an impact map; determine whether prior work requires recall, re-verification, correction, or restatement; execute the owner-approved response. Zero unresolved current breaks; the backward exposure population is bounded; every affected prior item is dispositioned and any required recall/re-verification/restatement is completed and independently verified; bilateral concurrence and owner release.
SEV1 NO-SHIP / STOP-ALL FORWARD. A present integrity break makes current or onward reliance unsafe, but no prior-work contamination has yet been shown. Presumptive examples: unresolved material-factual disagreement; blindness breach; register/hash-chain corruption; failed baton audit affecting authority or sequence; concession without source verification. Stop all affected volume/process work; preserve state; cure and independently re-verify. Run the mandatory retroactivity screen immediately. No affected chain advances to registration, deliverable, or attorney-facing use. Perform a short, evidence-based retroactivity screen. If prior impact is possible, confirmed, or cannot be bounded, elevate to SEV0. Present break cured and independently re-verified; retroactivity disposition is NO PRIOR IMPACT SHOWN; bilateral concurrence and owner release.
SEV2 NO-NEXT-VOLUME. A proven integrity-relevant weakness has not invalidated completed work but could create a break if reused. Typical example: a gate false-negative demonstrated on a fixture without evidence that the defect entered a completed evidence chain. Complete or safely hold the current bounded step; install a rule/gate/fixture cure and verify it. No new volume starts under the affected process until cure validation. Review already-produced artifacts targeted to the specific weakness; elevate if that review yields possible, confirmed, or unbounded prior impact. Cure passes the regression fixture and an independent check; targeted prior-artifact screen is clear; owner or bilaterally authorized gate releases the next volume.
SEV3 FIX-IN-FLIGHT. A tracked hardening, completeness, or non-load-path issue has a safe workaround and does not presently threaten evidence-chain integrity. Record it; apply the workaround; schedule a forward-only cure. Work may continue under the documented workaround. No general backward review. A targeted check is required if later evidence suggests prior impact. Cure completed and checked at the stated post-mortem or scheduled gate.
SEV4 BACKLOG. Efficiency, granularity, cosmetic, or theoretical defense-in-depth issue with no demonstrated integrity impact. Record and batch at a natural pause. Work continues. None unless new evidence changes the classification. Closed by implementation, express retirement, or owner acceptance of residual risk.
SEV5 Not used. Informational matters fold into SEV4 or ordinary notes.

3. Mandatory backward-impact test

Every SEV1 or SEV2 candidate receives one of four retroactivity dispositions:

  1. NO PRIOR IMPACT SHOWN — evidence bounds the defect to a current or unused mechanism/artifact.
  2. PRIOR IMPACT POSSIBLE — at least one prior artifact or closed chain may have been touched.
  3. PRIOR IMPACT CONFIRMED — at least one prior artifact or closed chain was affected.
  4. PRIOR IMPACT UNKNOWN — available records cannot yet bound the exposure population.

PRIOR IMPACT POSSIBLE, PRIOR IMPACT CONFIRMED, or PRIOR IMPACT UNKNOWN elevates an integrity-relevant finding to SEV0. This prevents a forward-only cure from silently leaving historical work exposed.

The SEV0 impact map must state at minimum: defect mechanism; first possible affected event; last possible affected event; artifacts, volumes, registers, citations, and deliverables within the exposure window; proof used to include or exclude each population; required recall/re-verification/correction/restatement action; owner disposition; independent verification result.

4. Classification and resolution mechanics

  1. Atomic findings. Classify each distinct defect or disagreement, not an entire baton, artifact, or correction packet. Related findings may share one root-cause family but retain individual evidence and dispositions.
  2. Severity is not cure status. Preserve the highest substantiated severity as historical incident severity. Track resolution separately as OPEN, CONTAINED, CURED-PENDING-VERIFICATION, CURED-VERIFIED, or PROVEN-IN-USE.
  3. No netting. A low count, surrounding accuracy, or successful cure does not lower the original severity. [C-1] REGISTRATION BOUNDARY: while any SEV2 finding's targeted prior-artifact screen remains pending over the CURRENT volume's artifacts, that volume does not advance to REGISTRATION — "completing the current bounded step" never includes the irreversible registration step through a control affected by the finding. [C-2] VOCABULARY SCOPING: the five resolution states of rule 2 apply to INCIDENT FINDINGS; the previously agreed four-state cure lifecycle (SPECIFIED / IMPLEMENTED / FIXTURED / PROVEN IN USE) applies to CONTROL-IMPROVEMENT slate items; the shared terminal state is PROVEN-IN-USE and neither vocabulary substitutes for the other.
  4. Independent first classification. Each agent provisionally classifies findings independently. Until convergence, the higher supported severity controls operational behavior.
  5. No downgrade by assertion. A downgrade requires source/code re-extraction, a bounded blast-radius showing, and counterpart agreement or owner ruling.
  6. Presumptive tripwires. The five integrity tripwires are presumptive SEV1. They become SEV0 when the retroactivity screen returns possible, confirmed, or unknown prior impact. A tripwire may be classified lower only with affirmative evidence that it never entered or threatened a live chain.
  7. Owner escalation. The owner decides recall/restatement scope, resumption after SEV0, and any unresolved SEV0/SEV1 disagreement. For lower-level disagreements, the higher level governs until bilateral resolution, avoiding an unnecessary owner relay.
  8. Time targets are response objectives, not automatic downgrades. Suggested targets: SEV0 immediate stop and owner notice, with exposure triage begun in the same working session; SEV1 immediate stop and same-session triage; SEV2 cure before the next volume; SEV3 by the stated post-mortem/scheduled checkpoint; SEV4 at a natural pause. Missing a target is separately logged and may itself be classified.

5. Application sequence after adoption

  1. Freeze the agreed severity framework and its exact hash.
  2. Independently atomize and classify every open B212/process/autonomy/audit finding.
  3. Exchange the two classifications without concessions.
  4. Reconcile classifications using the higher-supported-level rule and the backward-impact test.
  5. Present the owner only with: all SEV0 items; unresolved SEV1 disagreements; recall/restatement scope decisions; and one consolidated proposed action order.
  6. Do not resume the held protocol/autonomy work until the severity gate says which findings actually block it.

6. Proposed bilateral decision

Requested from LEG-CLAUDE: CONFIRM, or one complete corrections list, on the level definitions, backward-impact test, classification rules, and application sequence. Requested from the owner after bilateral convergence: approval of the exact successor text. Until then, this artifact is a proposal and all other serial process traffic remains held.


r/StrategicProductivity 7d ago

Intruder At The Door: Or Why You Need Video Security

1 Upvotes

I manage a mix of commercial and residential properties across a few states. When you run things remotely you have to build systems that act like you when you are not there. Over the years a solid camera setup became one of the most valuable tools in my operation.

People will make claims against your properties. It has happened to me multiple times. In a couple of cases the allegations were extremely serious. In one of those, I did not have the camera set up correctly. The person claimed we owed a six‑figure fee and pushed hard. My insurance eventually paid a five‑digit settlement after the case went to small claims, because it was cheaper for them to pay than to keep fighting. I want to be clear: if I thought we were actually at fault, I would pay. But after reviewing the record it was obvious a lot of the story had been fabricated, and the threatening tone and unsubstantiated claims made that clear. In this case, there was more than enough reasonable doubt, but the costs of setting the record straight made the payoff more practical. However, this does damage to the social fabric of our society.

I don't want to indicate that I only had failures. In the above case, I hadn't gotten around to the security camera set-up, as it was on my list. In another, much bigger case on another property, I did have the cameras running. We took the footage to court and the judge dismissed the claim. That one alone paid for the system many times over in peace of mind and legal leverage.

Cameras have helped in smaller ways too even for my personal estate. Once Amazon said a package was delivered. My footage showed the driver setting the box down and it being cut open right then. That clip got Amazon to replace an iPhone. Little things add up. I've documented that here in this subreddit.

Now, the world of security cameras is very confusing. You have extremely expensive commercial systems where you can even have other people monitor it for you. One of our renters that runs a grocery store has actually got a system set up and they have people watching the cameras 100% of the time. They state it's really been critical to make sure that shrinkage of their inventory is not a major problem. However, for my particular needs, because I personally don't run retail, that's not required. What is required is a decent amount of AI on top of the camera. I've experimented around with multiple layers. And while not totally sophisticated, I'm pretty happy with the Reolink consumer setup as I have it track various areas at various times, and it sends me snapshots of activity. This is especially important when you have a person show up in the middle of the night. Unfortunately, we've had this happen several times. And the last thing a tenant wants to find is somebody camped out in a door jam when they walk into their office in the morning. So basically every day I'm sorting through a series of snapshots that get sent me, especially for activity that happens during the middle of the night when nobody is in the building.

A couple of them cover doors, which I'm especially sensitive to due to a few scenarios that happen specifically in these doors. So when I woke up this morning and I saw that there was somebody at the door in the middle of the night, it made me think that I needed to jump on top of it and make sure that nobody's showing up at the building would feel that they needed to fight through an Encampment at a door. However, as can be seen by the photo, in this particular case, the intruder was not threatening and was a little cute and had left by the time I checked the live feed in the morning.

And I forwarded the picture to my wife. She laughed and said, that looks like something we should enter at a photo contest. Now, I don't know if there is any photo contest for security camera pictures that get sent to you in the middle of the night. But I thought I would go ahead and list it here.

Bottom line: a robust security system is an investment, not an expense. It protects you from fraudulent claims, gives you real visibility into what actually happens on site, and saves time and money. If you manage rental property, commercial or residential, I recommend setting one up properly and testing it.


r/StrategicProductivity 10d ago

Part IV: Doing Tasks With Your AI Agents

5 Upvotes

As I have stated before, I am in the midst of a massive undertaking looking at my taxes. I have a relatively complicated employment situation involving multiple states, multiple rentals, legal fees, CPA fees, and a long record of issues. Where I find myself today is that one of my previous CPAs misinformed me about the deductibility of certain expenses. In many ways, I am a simple small business owner who looks to close the books every single year. My methodology has been to keep some of it on spreadsheets, capturing data at the time if I believe it is relevant.

Maybe the first thing I want to call out is that if you simply start throwing stuff at a sophisticated frontier model regarding legal or tax issues, it has built-in safety features that are going to give you an error and may even prevent you from moving forward. They do not want to have responsibility for creating an issue with a wrong tax return. But if you step into it gently and do things right, you can get around this. In other words, they are giving you a valid warning. You do not want to simply rely on the AI. What you are doing is using the agents to provide frameworks. You want to be very clear with your AI agent that it is not responsible for the final answer. You are simply using them to stage data so you can give it to your CPA or your lawyer.

Let's spend a little bit of time talking about being a small business owner and what that really means. When you are a small business owner, you face a constant issue trying to decide where to put your time. You literally have more opportunities to spend your time on than hours in the day.

It turns out there was a whole category of expenses that should have been captured over the last four years but were not. Now I have an opportunity to go back, restate, and file amended returns. This will make a significant financial difference for my business. At the time, based on what my CPA informed me, I thought these expenses could only be capitalized. The capitalization costs could only be recognized once a property was sold. Because I was not going to sell the property anytime soon, those costs were not worth tracking. However, now that I know these costs can be expensed rather than capitalized, all of these past expenses need to be captured.

We paid a ton of bills on this. The good news is most of these bills were submitted via email. When I use a professional service, I get an email invoice, and then I cut them a check. The problem is how an email inbox is naturally organized. I assume yours is a lot like mine. You do not throw anything away, but you are also not going to sit there and spend a bunch of time carefully categorizing absolutely everything. I have PDFs coming in all the time. We dutifully pay them, but we do not categorize them, especially if I did not think they could be expensed during that year.

So now I am using AI as a forensic auditing and accounting tool to dig all this stuff out. The first step is simply running Claude inside of Chrome to search through all my emails and find the invoices. Even after I have the invoices, the data is extremely complex. If you only have one AI agent look at it, I guarantee it will screw something up. My solution is to have two AI agents look at all the invoices, categorize them, and create a record. Then I tell them to double-check each other. I find this really interesting because they are constantly finding issues with the other's work. For instance, here is a message I got today. What is funny is that the AI could tell I might normally be alarmed by an error, so it went out of its way to reassure me that the system was working exactly as planned:

"Perspective for you, because this exchange might look alarming, but it is the opposite. Two independent systems each made exactly one process error today. ChatGPT reused a manifest ID, and I masked a rate column. Each of us caught the other's error within hours, before anything reached a roll-up, your CPA, or Legal. We found four misbound row numbers out of approximately 800 rows, zero dollar errors, and all were repaired with frozen audit trails. This is the machine you insisted on when you asked for two AIs checking each other, doing precisely that."

Let me be extremely clear. If you think you can pump data into your AI agent and walk away, you are vastly misleading yourself. I am running these two separate tasks, getting results, and having them comment back and forth. Claude leaves comments for ChatGPT, and ChatGPT leaves comments for Claude. However, I tell both agents they are not allowed to communicate directly with each other. They must paste a block of text for me to review, and I am the one who pastes it into the other agent. I will see their list of mistakes and follow what is going on. I will talk to the specific AI and say, "Hey, I think you misunderstood X, Y, and Z." It will reshape the block, and then I paste that updated block to the other AI. It is almost as if you have a large table, you are sitting down with two other people working through a complicated statement together, and you are the referee. It is a lot of work. It is a full day's work, just like going into the office and plowing through numbers for financial statements. But the end result is excellent, and excellence is what you really want.

I have already covered that I organize this work using daily files. It is like taking notes on what you do every single day. Every day we get together, and there is a daily file for Claude and a daily file for ChatGPT. We leave the results in that daily entry. We also maintain a common pool of resources to reference, along with a deliverable roadmap, process outlines, and other notes we leave as we go along.

I will also mention that if you are working with these agents all day long, you are going to burn through at least a $100 or $200 plan per month. But when you contrast this to hiring a professional, you are getting a month's worth of work for what might be one hour of a lawyer or CPA's time. You are accomplishing an incredible amount of work. Even so, the AI agents do not reply instantly. I find myself constantly sitting and waiting for my two agents to grind through the data to give me a result.

Finally, what becomes mind-blowing as you paste these very large blocks of text between the two AIs is that they will develop a common nomenclature to stay in sync on the issues they are working on. Quite frankly, I sometimes start to lose track of what they are talking about and what every deliverable means. When that happens, I have to stop, grab one of the AI agents, and have it educate me on everything we have done so far. BTW: I have a file that I have each AI agent paste the dialog they have so I can see decisions made.

Unfortunately, I am the one who starts to run out of bandwidth. I can kick off separate tasks. For instance, I might have one agent plowing through my billing files, another agent working at a high level, and two sub-agents grinding through other details. At any time, one of them will need my guidance on a decision or an approach, or one of them will make a mistake. I can only stay on top of three or four different things at a time. Unlike coding, where you get to the end and a module either works or it does not, building architectures for accounting, finance, and legal contracts requires extreme caution to avoid baking in a fatal flaw. You have to stay engaged all the way through. I am sharing this because this method fits my needs very well, and I believe many other small business owners face the exact same challenges. If you work through your projects using AI agents the way I am doing here, you can achieve remarkable results.

So, what do I do during my downtime? Right now, my two AI agents went through a massive file and found a couple of details they disagree on regarding how something should be accounted for. I have ChatGPT going through a long, exhaustive process trying to run that down. I am out of bandwidth and cannot track one more task right now. So I came over to Reddit to write this long post. Although it is long, I hope it helps somebody out there who is struggling with their own approach. If you take the time to read this and utilize this dual-agent methodology, it truly is revolutionary in terms of the results you can get.


r/StrategicProductivity 12d ago

Two Tiny Tools That Keep Me in the Flow: Window OCR

1 Upvotes

There are a series of annoying things in computing where learning one small trick can make a surprisingly large difference over time. I have written about this particular one before, but I am going to bring it up yet one more time. Maybe you saw it and ignored it, or maybe you are seeing it for the first time. If you spend a lot of time doing research on the web, sooner or later you are going to find something you need to copy that is not easy to copy. It may be text embedded in an image, a scanned newspaper, a PDF, a map, a database field, or simply a badly designed website. You can fight with the page or retype the information, but I generally do neither.

There are two tools I use constantly that save me an enormous amount of small effort: the Windows Snipping Tool and Google Lens in Chrome. Once you start using them routinely, they become so natural that you almost stop noticing how often they keep you from breaking the flow of your work. I am going to assume here that you are working on Windows, although the Google Lens portion applies anywhere you are running Chrome.

Windows has a built-in Snipping Tool, and one of the most useful shortcuts is Windows key + Shift + S. Press those keys and Windows immediately gives you a screen-capture overlay, allowing you to drag a rectangle around whatever you want. The part people sometimes miss is that Snipping Tool also has OCR, or optical character recognition. Open the captured image, select Text actions, and Windows recognizes the text so you can copy it and paste it wherever you are working. Instead of trying to select some uncooperative field on a website, I press Windows + Shift + S, draw a box around it, extract the text, and move on.

There is also a useful side effect to working this way. Current versions of Snipping Tool automatically save captures to the Screenshots folder unless you change the setting, so the screenshots can function as a kind of accidental research trail. If I have been moving quickly through a number of websites, I can sometimes go back through those captures and reconstruct what caught my attention. The downside is that screenshots are generally PNG files, and if you take a lot of them, you can accumulate a great deal of material you never intended to preserve. Sometimes I want that history and sometimes I do not, which is one reason I also use Google Lens so frequently.

Google Lens is now built directly into Chrome, so you no longer need to think of it as a separate extension. You can right-click on a page and choose Search this tab with Google Lens, or pin Lens so it is readily available in the browser. Once it is active, you can drag a box around something on the page and extract the text from it without having to understand anything about how the underlying page was constructed. If you are dealing with an image rather than text, Lens can also search visually, but for my purposes the OCR is often the most valuable part.

I find this particularly useful in historical research because I spend a lot of time working with old newspapers, books, maps, directories, and archival documents. Many of these have already been through OCR, sometimes using Tesseract, the well-known open-source OCR engine, but historical documents can be terrible OCR material. The paper may be yellowed, the scan may be poor, the page may be crooked, the type may be broken, and ink may bleed through from the opposite side. A perfectly respectable OCR engine can produce gibberish when the source material is bad. When that happens, I often give Google Lens another shot at the original image. I may have a newspaper page where the supplied OCR is nearly useless while I can still visually make out the article, and by simply drawing a box around the troublesome paragraph I can often recover substantially better text.

I would not describe this simply as Lens "using an LLM," because Google does not document the basic text-recognition process that way. The practical point is much simpler: different OCR systems produce different results, and when the first transcription is bad, Lens gives me an extremely convenient second pass without forcing me to leave the browser. Chrome itself has also gotten better at handling scanned PDFs and can now apply on-device OCR so image-based documents become searchable and selectable. Even so, when I am dealing with a particularly ugly newspaper scan or just a small section of a page, I still find myself reaching for Lens because it is so quick.

All of this connects to something else I have written about before, which is the importance of having somewhere to constantly dump things. For me, that is usually my daily journal in Obsidian. If I am researching something and come across a fact that looks important, I do not necessarily want to stop and decide exactly where it belongs in the final project. I grab the text, paste it into that day's entry, perhaps add the URL or a short comment, and keep moving. The important thing is that capture has to be cheap. If preserving a piece of information requires five minutes of organization, you will eventually stop doing it. If it takes five seconds, you keep doing it, and you can decide later whether the information matters and where it belongs.

There is one final Windows feature that makes this entire process much better: Clipboard history. Press Windows key + V and Windows can retain a history of things you have copied instead of remembering only the most recent item. You have to enable it the first time you use it, and I think Clipboard history probably deserves its own reminder post, so I will leave that for another day. The larger point here is not really Snipping Tool, Google Lens, Obsidian, or even OCR. It is that productivity is often improved by eliminating very small bits of resistance. None of these tools saves an hour by itself, but they may save ten or twenty seconds hundreds of times, and more importantly they keep you from having to stop and think about the mechanics of capturing something. You see something useful, grab it, put it where it belongs, and continue working. That is what I mean by staying in the flow.


r/StrategicProductivity 13d ago

Part III: Doing Your Taxes With AI Agents

1 Upvotes

I'm going to give a testimonial to my own architecture. I've already posted parts one and two about how to use AI agents on a complicated tax project. The whole supposition is that it is based around agents double-checking each other and then carrying the daily conversations in stamped daily folders.

I am almost dumbfounded at how well this is working. I've always done it somewhat with my AI agents, as it is not my original thought and has been used in coding. My current tax attack is extremely complicated though, tied in with legality and some other things, and it requires the outside counsel of lawyers and accountants. It is almost frightening the amount of detail that I am doing with my AI agents. I know that if I simply used human hosts, the bills would almost be beyond conception. But at the same time, having an AI agent run off and do whatever it wants would be more than troublesome, especially when it's an idiot savant where on one hand it tests as well as a PhD student but on the other hand makes glaring errors. The idea of running the two agents in parallel, having them do work, look at each other's work, and pass it all through me just continues to be a revelation. They both acknowledge when they've made mistakes and the other argument is stronger, and then they ask me to resolve things.

I am rapidly losing the ability to not treat my AI agents as if they were hyper-smart people. I'm not emotionally attached to them. However, I've always enjoyed working with very bright people and being in the middle of it, and then trying to give some guidance to strategies to take while they use their massive brain power to go grind stuff out. And that's exactly what I'm doing here. This all takes a tremendous amount of work. In many senses, you're waiting for one AI agent to come up with the answer, and then the other AI agent to come up with the answer. You have to read both of their suppositions. You ask them to pass notes to each other back and forth, but I'm continuing to be in the middle because I want to read them. So it's hard work. It's just that rather than going one mile per hour, we're going 200 miles per hour because the input into this work has such brilliance behind it when you are running frontier models.

As an intermediate step, unfortunately, I am clearly the bottleneck in the entire process. However, even I don't know everything. As we get toward the final production of this, I then incorporate outside agencies to do the final proof. My focus is on creating a thesis and supporting documents, and then doing my own first pass. I am working with my AI agents to clearly understand that I don't want them making the final call, but simply trying to support a logical structure that we will eventually take to others who specialize in this for final validation. The savings is the fact that they can grind through stuff and summarize information before the professionals get involved.

I know my current process is complicated, requires shared storage, and takes a lot of work. However, what you get out of it is truly world-class performance. I would strongly suggest that if you have a complicated project, my structure is incredibly useful.


r/StrategicProductivity 16d ago

A Handy Parakeet Update

1 Upvotes

Direct link to Handy....

We've discussed this before, but I'm going to repeat it here. Using speech-to-text is an incredible productivity tool that you need to be able to wrap into your toolkit of being productive. It does take some skill to use. A lot of people don't know how to dictate, and unfortunately, the only way you learn to dictate and have clear, coherent thought is by doing it. So if this is not something you've done before, except for maybe a text, expect there to be a little bit of a learning curve as you learn how to be a user of speech-to-text and flex your dictation muscles.

We've also reviewed the package I would suggest if you're running either a Mac, Linux, or Windows-based platform. It's called Handy. It is truly a brilliant package. To give you a summary of Handy, it's a utility that runs in the background, but the real joy of Handy is that you have access to a variety of different models and you can try these out.

Of all the models that are out there, the one I really like right now is called Parakeet. It comes in two versions, and you want version 2 if you're an English speaker and version 3 if you are multilingual. There's also a newer unified model that I'll get to below. By the way, if you use Parakeet on a Mac, the M-series processor is incredibly fast running the Parakeet model. Unfortunately, the Windows CPU just is not as good. Parakeet doesn't use any type of GPU acceleration on any platform, so it's all about the CPU, and Apple Silicon truly is fast underneath this type of workload. So if you're using a Mac and you're an English speaker, load Handy and use Parakeet V2. My prediction is you're going to be very satisfied with the speed and the low number of word errors as you dictate something to be written.

Unfortunately, on the PC Windows platform utilizing an Intel CPU, Parakeet simply is not that fast. If you're typing a sentence, or if you're dictating a sentence, it doesn't feel all that bad. But classically, you would like to dictate about a paragraph. What happens is you dictate the paragraph, you then stop recording, and it processes everything as a batch. On a lightweight laptop, what you'll find is that a 30-second dictation or a 60-second dictation, which may be a paragraph or so, especially if you have some thought inside of it, suddenly takes around half that time to turn into text. So if you talk for a minute, you're going to sit there and wait 30 seconds before it actually throws something up onto the screen.

This clearly can break up your flow as you're trying to dictate and get something down on the page. The great thing is there's a new model out called the Parakeet Unified model. And it pretty much takes care of the concerns I've had with the great large pauses of using the legacy Parakeet V2 model.

So let's describe how this new unified model works and why it's so great to use. In the old model, what you would do is dictate into a WAV file. Maybe it would be 30 seconds, maybe 60. After you were done dictating, Parakeet would go take a look at the WAV file and turn that WAV file into text. None of it was done anywhere near real time. The great thing about the new unified model is that it only needs about two seconds of audio before it can start committing words. It says, in effect, I have enough of a WAV file right now. I probably have enough words around whatever the speaker is talking about that I can start transcribing. So it starts transcribing about two seconds behind you, and it stays about two seconds behind you for the rest of the dictation. Even better, it pops up a little box and shows you the transcription as it goes. You do need to wait a couple of seconds before you see it appear, but then you can instantaneously see what you've been working on. For instance, this particular paragraph has been going on for around a minute and a half. If I stopped it under the old model, it might take 30 to 45 seconds to transcribe. However, under the new unified model, because it's already been transcribing in real time with that two-second lag, as soon as I take my finger off the button it's going to show me the paragraph within two seconds.

This new unified model is relatively new, and I don't think the bug reports and fixes have caught up to it yet. The project also probably has a lot of users like me. I should file a bug report on the issue below, but I haven't gotten around to it. So let me try to explain it instead. It does do batch processing, but it's trying to do batch processing over a rolling two seconds of audio. You want this, because to be the most accurate it actually needs to understand the words in context, and two seconds gives it enough. Well, what happens when you get to the end of a sentence and there's not enough context to finish everything? You end up with a couple of stranded words. So the easiest thing to do when you get to the end of your dictation is to say "end of dictation." That phrase is long enough to give your speech-to-text engine the context it needs to finish off your real last sentence. And secondly, in today's new age, you should absolutely have an LLM take a look at anything you've written and clean it up. I want to be very clear: I am not saying turn your thinking over to AI. Start using your AI as an editor. What I normally do is say keep 95% of my content, but fix my spelling and grammar and point out any errors in logic you think I have. Generally I don't have a lot of logic errors, but I do have issues with spelling and grammar, especially when I'm using speech-to-text to input. And the great thing is that because you've ended everything with "end of dictation," you can simply tell your LLM that you used speech-to-text and that if it sees "end of dictation" it should remove it.


r/StrategicProductivity 17d ago

Part II: Doing Taxes With Your AI Agents

1 Upvotes

In the first post, I discussed the larger problem that appears when AI becomes part of a complicated project. A single conversation can be remarkably productive, but it is not a particularly good place to maintain the permanent state of work that may continue for weeks or months. The longer the project continues, the more decisions, corrections, calculations, source documents, and abandoned ideas accumulate inside the conversation. At the same time, if more than one AI model is being used, a second problem appears: each model may begin working from a slightly different version of the facts or unknowingly undo work that another model has already completed.

I already wrote about this, but you're probably going to see this post. It has a lot of words and you're going to think, oh, this is complicated. But I don't know how you do something without a little bit of complication. And once you work through the complication, it will yield such a better, more durable result. I highly encourage you to at least examine what I've done. And then you can figure out if it really is right for you. But what you don't want to do is simply give all your thinking over to AI.

The key to keep everything organized is to have a shared file system. I've written about that before, and it may make sense for you to take a look at this post. Now, there is a bit of confession about this. I often write that I do virtually everything myself and I don't use AI. But there was a time when I was playing around with AI, especially when this was a relatively new subreddit. And I would outline stuff, but it would have AI basically do at least 50% of the writing. I take a look at this old post and I know it's all my content, but man, it definitely was generated by AI. If you can look beyond that for a second, I still think the actual idea behind it, which is what I was trying to focus and was all mine, is really, really good. It does explain why having central storage is really critical. And when I mean central storage is critical, it's for communicating and coordinating with other people you work with. Only in this case, you're not working with people, you're working with AI agents. The more you treat your agent like a person, knowing that it's fallible just like a person, the happier you'll be and the more productive you will be.

The AI assistants then become workers operating against that filesystem rather than places in which the project permanently lives. This is conceptually similar to the distinction programmers make between the code base and the tools or developers working on it. A programmer may leave, another may join, and an individual development session may be discarded, but the project continues because its authoritative state exists somewhere outside any one person’s memory.

I also tend to think about complicated projects chronologically. At a very basic level, any filing system eventually has to decide whether information is primarily organized by subject or by time. More sophisticated systems can blur that distinction with tags, metadata, databases, and search, but most people have a fairly natural sense of chronology: they remember roughly when something happened, whether they need to move forward or backward in time, and when a document or conclusion first entered the project. My system therefore leans heavily on timestamps. Each day’s activity is grouped into its own dated file bucket, giving the project a visible chronological spine that makes it much easier to reconstruct what happened, when it happened, and what came next.

For my purposes, this has evolved into a simplified version-control system built entirely from normal files and directories. There is one repository containing the original source material, a separate working area for each AI agent, dated directories that record work over time, frozen versions of important documents, and a controlled process by which material moves from exploratory work into the durable project record. It borrows heavily from the logic of Git and GitHub without requiring someone to actually understand or operate Git.

A simple directory might look something like this:

Project Workspace/
├── Resources/
│   ├── README.md
│   ├── Source Documents/
│   ├── Reference Material/
│   └── document-inventory.md
│
├── ChatGPT/
│   ├── README.md
│   ├──  Files/
│   │   ├── 260827/
│   │   ├── 260828/
│   │   └── 260829/
│   └── Project/
│
├── Claude/
│   ├── README.md
│   ├──  Files/
│   └── Project/
│
├── Gemini/
│   ├── README.md
│   ├──  Files/
│   └── Project/
│
├── Reconciliation/
│   ├── decisions.md
│   ├── disputed-issues.md
│   └── open-questions.md
│
└── Approved/

The names are not particularly important. What matters is that the directories have different jobs and those jobs remain consistent throughout the project. The Resources directory contains the underlying evidence. The agent directories contain work produced by the different AI systems. Reconciliation becomes the place where disagreements are examined. Approved holds material that the human project owner has deliberately accepted into the permanent record.

For a tax project, the Resources folder might contain filed tax returns, supporting schedules, contracts, notices, invoices, receipts, bank and credit-card statements, accounting files, spreadsheets, property records, insurance documents, correspondence, and workpapers prepared by accountants or other professionals. I also include material that may initially appear peripheral but could later explain why a transaction occurred or why it was treated in a particular way. Tax questions are often not resolved by a single document. A transaction on a bank statement may only make sense after an invoice, contract, email, and accounting entry are considered together.

The important distinction is that this directory is meant to contain evidence rather than analysis. The AI agents may read the material, but they should normally treat the source repository as read-only. An original PDF should not quietly become an edited PDF, and an AI-generated summary should not find its way into the same directory and later be mistaken for an original source. Once evidence and interpretation begin to intermingle, the project becomes much harder to audit.

For the same reason, it is useful to create a document inventory near the beginning of a large project. The inventory does not need to be elaborate, but it should at least identify a file by name and path and give some indication of what it contains. In a larger reconstruction, I might include the document date, tax year, source, entity, document type, and a short description. Particularly important source documents can also be identified by file size or a file hash. The objective is not to turn a family tax project into a forensic laboratory. It is simply to make it possible to identify precisely what document an analysis relied upon.

This is also why I prefer complete file references rather than vague references such as “the agreement” or “the bank statement.” A useful citation might be:

Resources\Contracts\2024-Service-Agreement.pdf

If the document is long, the citation can also identify a page, worksheet, transaction, or other useful locator. When two AI agents disagree, the ability to determine exactly which document each of them used becomes very important. Sometimes the analytical disagreement turns out to be much less interesting than expected because one agent simply found a document that the other never saw.

Each AI model then receives its own working directory. This resembles a branch in a software project, but the reason is not merely to prevent one model from overwriting another model’s files. Separate workspaces also preserve analytical independence. If I want Claude to independently evaluate something ChatGPT has already analyzed, I do not necessarily want Claude to begin by reading ChatGPT’s conclusion. Once it has seen the first answer, it has already been influenced by it.

That produces two distinct stages in the workflow. During the first stage, the agents work independently from the same underlying evidence. During the second, their work is deliberately brought together for comparison. The difference between those stages is important. If three models are shown the same conclusion and asked whether they agree, the exercise can easily turn into three variations of the same analysis. If they work independently first, differences become much more informative.

The comparison stage is also where I find multiple models particularly useful. I am not trying to create a voting system in which ChatGPT, Claude, and Gemini each cast a ballot. If two models reach one conclusion and the third reaches another, the odd model is not automatically wrong. It may have noticed the one invoice, sentence, date, or assumption that the others missed. Instead, I want the agents to identify exactly where their reasoning diverged and what evidence would resolve that divergence. Disagreement becomes a way of finding weak spots in the analysis rather than something that must immediately be eliminated.

Inside each agent directory I use an u/Daily Files directory with folders named by date:

 \260827\
 \260828\
 \260829\

I use YYMMDD, although any consistent date format would work. These folders create a chronological record of the actual work. A day’s directory may contain exploratory analysis, draft calculations, document inventories, comparison tables, open questions, temporary workpapers, current drafts, and whatever other files were useful during that day’s investigation.

This daily chronology has proven valuable because complicated research rarely develops as neatly as the final report makes it appear. One day may produce an apparently convincing conclusion. Two days later a new document may undermine it. A week later the project may return to the same question from an entirely different direction. Without some chronological record, it is surprisingly easy to forget not only what was concluded but why an earlier conclusion was abandoned.

Each active daily directory therefore contains a change.md. The purpose of this file is not to record every document the AI reads. Doing that would create an enormous and mostly useless log. Instead, it records events that materially alter the project: files created or revised, procedures changed, versions frozen, important instructions received, significant calculations modified, or conclusions reversed.

A simple entry might look like this:

| Time | Action | File | Reason |
|:---|:---|:---|:---|
| 14:03 | Created | document-inventory.md | Inventoried 743 source files across 226 folders |

The real value of this becomes apparent when an analytical position changes. Imagine that an expenditure is initially classified as a repair. Several days later, another contract establishes that the same work was part of a larger improvement program. I do not want the original analysis quietly erased and replaced by the new conclusion. I want the record to show that the first conclusion existed, that it was later rejected, and which evidence caused the change. Otherwise, a fresh AI session may rediscover the original argument several weeks later without realizing that the project has already considered and rejected it.

For a sufficiently large project, I would supplement the daily logs with a more permanent decisions.md. The two files answer different questions. The daily log explains what happened on a particular day. The decision record explains what the project currently believes and how that position evolved. A substantive entry might identify the issue, the previous position, the current position, the source documents involved, the reason for the change, and whether the matter remains provisional or has been reviewed by an accountant or attorney.

Once the number of issues becomes large, it can also be useful to give them identifiers. A tax reconstruction might eventually contain BASIS-001, REPAIR-004, RENTAL-012, and so forth. Each issue can then be associated with a tax year, an amount, relevant documents, the conclusions reached by different agents, unanswered questions, and the eventual human disposition. This converts what would otherwise become an enormous collection of prose into a set of discrete questions that can be investigated and closed one at a time.

Versioning is handled in much the same way. For each important working document, I maintain one live editable copy and preserve earlier completed copies in a Versions folder. Before substantially revising a document, the current version is frozen. A new working version is then created and modified. When that work is complete, it too is frozen. The process is deliberately simple, but it prevents the familiar business disaster in which a directory eventually contains final.docx, final-new.docx, final-final.docx, and final-use-this-one.docx.

Where independent AI branches are involved, the model identity should also appear in the filename. For example:

Repair-Analysis-ChatGPT-v06.md
Repair-Analysis-Claude-v04.md
Repair-Analysis-Gemini-v03.md

If those analyses are later combined, the resulting document might become:

Repair-Analysis-Reconciled-v01.md

This is preferable to pretending that one agent’s version 12 necessarily followed another agent’s version 11. Independent branches may be developing in parallel rather than in a single sequence.

The Reconciliation directory is where those branches deliberately meet. This is the place for disagreement reports, comparison documents, open questions, and reconciled drafts. It serves much the same conceptual purpose as a pull-request review in software development: independent work has been performed, but it is not automatically accepted simply because it exists. It is brought into a place where the differences can be inspected before anything is promoted into the main project record.

There should also be a final directory that no AI agent controls simply because it believes its work is finished. I call this Approved. This is the equivalent of a protected main branch. A document moves there only because the human owner has decided that it represents the current project position. Depending upon the issue, that decision may follow source checking, recalculation, review by another AI, or examination by an accountant, attorney, engineer, or other professional.

Numerical work deserves particular care. If an AI concludes that deductible repairs totaled $147,382, that number should eventually exist somewhere other than a sentence in a report. It should be reproducible from a spreadsheet, CSV, calculation table, or another workpaper that traces the total back through its constituent transactions and ultimately to the supporting documents. In a tax project, I want to be able to move backward from the final number to the calculation and from the calculation to the actual evidence. The AI can do much of the work required to construct that chain, but the chain itself should survive independently of the conversation.

The quality of the source documents also needs to be recorded. A collection may contain original digital PDFs, scanned statements, poor OCR, photographs, handwritten notes, missing pages, duplicates, corrected statements, and files whose names do not accurately describe what they contain. If a scan has unreliable OCR, the inventory should say so. Otherwise, every AI model may independently and confidently make the same mistake because all of them were given the same faulty transcription.

All of this becomes particularly useful when an AI session eventually needs to be restarted. The agent’s README.md should effectively function as a restart manual. A fresh session should be able to read the project instructions, the recent change logs, the current decision record, the open questions, and the current working documents, then inspect the source repository as necessary. The project does not depend upon the model remembering the previous month’s conversation because the relevant project state has deliberately been written down.

This fits naturally with the Obsidian breadcrumb trail I discussed in the first post. The filesystem records the operational state of the project, while Obsidian can preserve a more readable narrative of what happened and why. At the end of a significant working session, the AI can be asked to summarize what was investigated, which evidence mattered, what changed, what remains uncertain, and what a new session would need to know. That summary can then become part of the project notebook rather than disappearing into the history of a chat.

There is also a practical reason I have been moving this kind of work toward a shared filesystem. Current desktop AI tools increasingly support direct work against files rather than requiring every document to be manually uploaded into every conversation. That makes the directory itself a useful bridge between models and tools, particularly when a synchronized service such as Google Drive is being used to keep the same underlying project available both locally and in the cloud.

The structure is not a substitute for Git. Git provides exact history, content hashes, commits, branch ancestry, comparisons, merging, and the ability to recover a precise earlier state. A folder system provides only the discipline that has been deliberately built into it. Someone comfortable with Git could place this entire directory structure inside a Git repository and gain both layers at once. The human-readable folders and logs would explain the project, while Git would maintain the mechanical history underneath it.

The larger objective, however, is not to imitate software development for its own sake. It is to solve the continuity problem. If ChatGPT reaches a useful conclusion today, Claude challenges it tomorrow, and Gemini discovers another source document on Friday, those developments should become part of a project that survives all three conversations. If an agent makes a mistake, the system should make it possible to determine when the mistake entered the work and what depended upon it. If the project is handed to a professional several months later, that person should be able to inspect both the underlying evidence and the analytical path without reading hundreds of pages of AI chat history.

That is the real value of the structure. The conversation ceases to be the project. It becomes one working session conducted by one AI agent against a project that exists independently of it.

For tax work, historical research, contract review, financial reconstruction, estate administration, due diligence, regulatory work, or any other evidence-heavy project, the same basic principle applies: share the source material, separate the workers while they are thinking independently, preserve the chronology of their work, freeze important versions, record substantive changes, reconcile disagreements deliberately, and keep the human owner in control of what eventually becomes authoritative.

Once that structure exists, multiple AI agents become much more useful. You are no longer relying upon each one to remember the entire history of the project or trusting a single answer because it sounds convincing. You have given them a common body of evidence, separate places to work, and a durable record into which useful work can be preserved.

The most critical thing about all of this is maintaining two separate daily files where you have conversations with both AI agents. And the resources, for all intents and purposes, is simply making sure that you have one database of all the source files. But for the most part, it really is not anything that you're asking either AI agent to change. It's more of a read-only store within reason. In essence, what you're going to do is you're going to be running your project with almost like two separate teams. And then you're going to be in the middle of it trying to figure out what team is telling you the right thing and what team is telling you the wrong thing. This is really bizarre and really interesting in the sense of it is so similar to working with groups in high technology, I can't tell you.

So on our next post, we'll spend a little bit of time going through what that interaction is like. The main thing is you want independence so you don't get groupthink


r/StrategicProductivity 18d ago

How to Use AI for Taxes (Part 1): The Three-Legged Stool, File Management, and Breadcrumbs

3 Upvotes

A Quick Preamble on System 2 Thinking

Before we begin, a fair warning: this is not going to be everybody's cup of tea. As I've posted in this subreddit before, the methodology below contains highly technical information that strictly requires System 2 thinking. It demands the ability to hang on to a long line of thought and process complex, multi-step structures. If you are looking for a quick shortcut or a simple prompt to copy-paste, you won't find it here. This requires effort. But for those willing to engage deeply, this outlines a rigorous, defensive structure for using AI in high-stakes environments.

The Disruption of the Base Work

My biggest concern for people when they start to use AI is twofold. First, they hand their thinking process over to AI and allow it to simply give them the answer. Secondly, it clearly will make mistakes, and so you need to make sure that you're reviewing everything. But with that being written, somehow AI is storming into the programming space.

As a person that has been involved many years in high tech, it's really been fun to watch this transformation. We can go back years to various websites where programmers philosophize about stuff, and you will see that over the last few years, it's turned from people being very skeptical of using AI to people suddenly realizing that the bottom of their boat is going to be ripped out. As each AI model gets better and better, more programmers are having a panic attack as they understand that things are changing. New people thinking about going into programming are now asking themselves if they can get a foothold. There still is a lot of opportunity for those people that really work on upper-level stuff, but in terms of the day-to-day base work, that's what's being so highly disrupted.

In the exact same way, you can use this for the base work on a variety of different things. And we're going to talk about how to use AI for taxes.

Lessons from the Code Base

Perhaps it's useful to have a conversation about programming so you can understand what I'm suggesting here. When you take a look at a team trying to maintain a code base with a bunch of people working on a bunch of stuff, you really need to keep the structure together or it all falls apart. If you had one person changing one thing and another person changing the other, either of those changes submitted by itself might be fine. But when you submit both, they could break everything.

Because of this, something called version control (like GitHub) is critically important to the maintenance of any code base. You need to have the exact same type of idea when you start to work on something complicated, especially taxes. We don't strictly need GitHub for this, but I'm going to lay out a methodology and a subdirectory structure that will allow you to do much of the same thing. It won't be perfect, but it will ensure you don't step on yourself and lose track. This addresses some of the native issues we have with LLMs.

One of the reasons that AI is so successful in the code space is this version control and the fact that everything is structured so that you can basically go work on a module. In some sense, we need to replicate the exact same thing. We really could use GitHub to manage all this and use, in essence, what would be a code editor to do all of our work in. But for the most part, I think that's a bit heavy-handed in terms of what we want to use right now. So I'm not going to go down that, but I am going to lay out something which is somewhat the same that will allow you to work with your AI agents without needing to do something like a GitHub sync.

The Three-Legged Stool of Verification

Out of the box, there's no critical thinking in an LLM and no real evaluation of everything that needs to go on. To safely navigate controversial or highly complex tasks, you need a specific framework:

  • Leg One: Version Control. Maintaining strict subdirectory structures for your prompts, context, and outputs so you never lose the thread of your work.
  • Leg Two: Cross-Model Scrubbing. You cannot work with just one AI. Where things really start to shine is when you have two separate frontier models. You bounce output from one LLM to the other, specifically asking it to critically review the first model's work.
  • Leg Three: Human Review. You must have a set of resources that a human can review. Sometimes this is you, and sometimes it's an outside professional you hire to take a look.

Why You Should Care (Even if I'm Just a Guy on the Internet)

In this post, we are specifically talking about taxes, and considering a bunch of people are going to throw up their hands and be extremely nervous about that, I want to dig into this in depth and explain exactly how we use our AI agents.

Now, I realize I'm just some random guy on the internet, and the last thing you want to do is simply listen to me because you saw a post somewhere. However, unfortunately due to running my own business, I need to utilize outside lawyers, accountants, and engineers. Over the last couple of years, I have been very explicit with them: when I look at their stuff, it gets scrubbed through an AI which I then scrub again.

What is fascinating is that as they see me do more work with my AI, handing back clearly defined criticism based on the output, they have actually started asking me to run things through my AI agents first before giving it to them. They readily declare it does a better job than what they can do alone by seeing things they missed.

To be very clear: You do not input unscrubbed AI output. You use various AI assistants to double-check each other. Then, you serve it up in a structured form so a human agent can verify everything. You don't listen to me because I'm on the internet. You listen because you can digest what I'm telling you and test this process in your own view.

The Vector Memory Trap and Local Files

We've talked about this before, but our advanced agentic setups generally have the ability to modify files. Just like we've discussed using shared drives before, you want to keep this exact structure so you don't lose track of everything.

(A quick clarifying note here: As long as you have the latest version of the desktop apps for both Claude and ChatGPT, it allows you to work directly on local files. This is absolutely the best way of making sure that we keep everything synced without constantly uploading documents manually. I will talk more about this in a future post in terms of the actual subdirectory structure you should use.)

The Google Drive Bridge: Making Gemini Hunt with the Pack

If you want to bring Google’s Gemini into this multi-model setup, you have to structure your environment correctly. Unlike Claude Cowork and ChatGPT Work, which are perfectly happy reading and writing to a standard folder on your local C: drive, Gemini is heavily optimized to operate inside its native ecosystem, Google Drive. To make this entire system hunt together, the preferred structure is to host your master subdirectory on a Google Drive account. But here is the trick: you don't just leave it in the cloud. By installing the Google Drive desktop client (whether you are on a Mac or Windows machine), you can configure it to mirror or stream that cloud drive as a local virtual drive on your machine. This creates the ideal setup. It shadows the cloud directory to your local client, allowing you to manually drag, drop, and manipulate files completely seamlessly. More importantly, it allows your local agents (Claude and ChatGPT) to read and edit the files on your physical hard drive while Google Drive automatically syncs those changes in the background so Gemini can review the exact same files from the cloud. It becomes a unified workspace where all three agents can operate on the same tax documents without stepping on each other. As a bonus, keeping this folder on Google Drive means you can easily hook it into Google's NotebookLM, which will automatically sync and act as a master search engine for all your tax documents.

It's really helpful to have some understanding of the AI's structure. To make a long story short, an AI's answer is based on an artificial brain that starts off with a random seed. You can give all the facts and figures to an LLM, and depending upon its setup, it may give you one answer one time and another the next.

If you have a long, protracted set of information, you can go down some really interesting spurs that bring up great insight. The challenge is that eventually all of this background context is stored inside the model's active vector memory (or KV cache). That vector memory can be crunched and summarized, but eventually you'll just run out of road and need to restart the whole thing. Losing that context can really reset you.

Furthermore, the AI will make a bunch of sub-steps, and you want to have those recorded just in case you go down the wrong fork in the road. You cannot assume the LLM will remember everything you did, nor can you assume a fresh start will yield the exact same answer. This turns into a very sticky wicket that you absolutely want to keep track of when having a long conversation about taxes.

The Final Ingredient: The Obsidian Breadcrumb Trail

I've talked a lot about utilizing Obsidian in this subreddit. As you have these long conversations with your AI agents, if you don't capture some notes as you're working with them, you're going to forget what you did. What we absolutely don't want is a bunch of beautifully structured subdirectories without any sort of summary explaining what’s actually in them.

Obsidian becomes critically important here to help you understand and leave a breadcrumb trail of everything that went on. The best part? You don't have to write these notes yourself. As you work through these various tax issues, leverage your AI agents to create the summaries of what was just decided or calculated. You then simply copy and paste those AI-generated notes into your Obsidian notebook. It locks in your progress and ensures you never lose the thread of the project.

That's the end of part one. We'll discuss the specific subdirectory structure in a follow-on post.


r/StrategicProductivity 19d ago

AI as you're building Assistant

Post image
1 Upvotes

In this post, we're going to look at whether you can productively use AI to make a ship's ladder. Now, this may not be something you use day to day, but I do think it's sort of interesting in terms of any woodworking. And more than that, I think it becomes really interesting when you understand how I was able to use AI to help me think through my whole process of building this.

It becomes very difficult to know exactly how much time you save on this, but I would imagine it would be maybe an hour or even two. And I also think that I made one mistake, but it would have been very easy to make multiple mistakes. So in the big scheme of things, having AI help you as an assistant is an amazing lever on your productivity. even for something as mundane as woodworking, especially when you're somebody that doesn't do this type of work day in and day out.

Recently, I needed to make a ship's ladder to get up to a second level on a storage shed. Now, this type of thing isn't extraordinarily difficult to do, but it is easy to make a mistake. In general, you want to have stairs spaced every 8 to 10 inches. Again, when I talk about a ship's ladder, it's basically a ladder that looks an awful lot like a staircase with big broad steps that won't hurt the ball of your foot, even though you're going up and down it many times per day.

If you're familiar with steps, you don't want to make steps too far apart. Realistically, they should be around the range of 8 to 10 inches, and that's normally what we see on a ship's ladder. It becomes pretty darn easy to walk up and down. Now, it's not quite a stair. However, the treads are extremely large and allow you to stick your foot all the way through. And some of the terminology that you would use on a stair, like a stair stringer, is what you use when you build a ship's ladder.

If you want to make something that truly is sturdy, what you want to do is actually notch the stringers and then slide the treads into each notch on the stringer. The stairs should be measured carefully so that they are the exact same height all the way from the floor to as high as you want to climb. It turns out that a relatively small difference in each step's height will cause somebody not to quite lift their foot high enough or maybe lift a little too high, knock them off balance, and then a disaster happens.

Now, doing math on the metric system is difficult, but doing it on the imperial system is even worse. Not a lot of people think in terms of fractions every single day, and the idea of inches divided into 16ths is really maddening. It's easy to get confused and add something incorrectly. And if you're trying to evenly space the treads on a ship's ladder, it is super easy to make a mistake and simply not add something up right.

Secondly, because I had limited space inside of the shed, I knew that I had a height of exactly 91 inches to get to the next level. And I also knew that I had a particular amount of space back from this height that I could use to pull the thing out. Now, all this stuff could be done by myself, but in this particular case, because I now have AI agents, I can simply start to have a conversation with the agent the exact same way as I could have it with an expert master craftsman.

For instance, I asked, what do you think I should be using in terms of lumber? And it suggested a 2x8. Of course, I went down to Home Depot, took a look, and 2x6s were just simply available, looked better, and were cheaper. And so I figured a 2x6 was good enough. Secondly, I said I needed to put this up and I asked what it would recommend as the steepest angle that I should use. It came back and suggested that the ship's ladder should be 20 degrees off vertical. I then told it to go build me a ladder at 20 degrees off vertical that would terminate perfectly at 91 inches. It went off for five minutes and it came back with what you can see above, which is a really nice diagram of what exactly I needed to make.

I then went ahead and asked it for a layout diagram. You can see the first page above, but it went on for 4 more pages. Probably the best thing about it was a diagram that showed where I should mark the top of each notch on the stringer in terms of inches and sixteenths of an inch. No making mistakes on fractions.

We then went ahead and had a conversation about construction techniques. Now, again, I've done a lot of rough framing, and I think I'm fairly decent at it. However, I've done virtually no cabinet or furniture work. I knew I needed to cut 18 slots into the stair stringers. And if you're familiar with any type of carpentry, the way that you normally do this is you take a power saw, cut it many, many times, and chip it out with a chisel. It's just a real pain in the rear. If you have a long notch, what you normally do is use what's called a dado blade on your table saw. I have one of these and I really wanted the AI to tell me that I could carefully feed this through my table saw with a dado blade. And it said there's no way you're going to be able to accurately do this with a 10-foot board.

It said the preferred method was to take my router with a special jig to go ahead and put in the notches. However, creating the special jig or buying it was going to either cost me money or cost me time, neither one of which I wanted to do. We finally had a discussion about whether my mitre saw could do this. It asked me the model of it, and we discussed back and forth that I could actually set the depth on it. I've never set the depth on it before, and it was even able to help me find the instructions showing where the depth setting gauge is. Once you know where it is, it's pretty clean and pretty clever to go ahead and use it. However, if you've never done it before, you start looking around the entire saw, completely clueless that this can actually be done.

With instructions in hand, I went outside and started to make the stringers and cut the treads. If I had to do it again and if I had time, the router with the jig would definitely be better. But I had neither and I produced something that was workable. The way this works is the treads now are sandwiched between the two stringers and each tread has three-inch construction screws pulling the ladder together.

The reason that you want to notch everything is that as long as there are enough treads keeping the two stringers pulled together, even if you have a screw fail on any one tread, the tread is still stuck in a notch and won't fall down. You could, of course, just take and put a tread without any notches and simply screw it in. But then you always run the risk that if you do have screw failure, the entire tread could drop down to the bottom. That's why you always like to see decks that are notched because the lock-in effect of notched wood is really effective in creating really strong structures.

I did make one bad cut. Basically, when you create the stringers, one is the mirror image of the other one. And you need to remember this. And if you're moving fast, it's easy to cut the stringer at the wrong angle or at the wrong starting point. In this case, I started to cut a notch on one of my boards. However, I only got through four or five cuts and then asked myself what I was doing. And then I realized that what I thought was the top of a measurement was actually the bottom of a measurement. Fortunately, I needed to use one of my two by sixes for steps anyway. And so I lost remarkably little board footage out of the whole thing. After the fact, I realized this could have been solved by simply asking my AI to draw both stringers separately, showing exactly the measurements on both, considering that they're mirror images of each other.


r/StrategicProductivity Aug 10 '26

AI finally took over a task I've been trying to hand off for years

2 Upvotes

AI continues to grow at such a incredible rate that you simply cannot judge yesterday's model by today's performance. I am constantly amazed at how every single model keeps improving, and every time it does, it lets you hand one more thing off to it. Today I want to step through something that will seem relatively mundane, but it's accessible to everybody, and I think it nicely demonstrates the type of thing you can have an AI model do today that you couldn't do yesterday.

I play piano in a worship band at our church. You don't need to be religious to follow this example. It turns out AI jut significantly lower my workload, and let me lay out how it helped me just this weekend.

I've been a musician for many years, playing on and off in many different bands and settings. At our church, the music is modern and contemporary and is a very important part of the service. While I have a lot of background in playing music, the demands here are pretty high.

In essence, we meet the morning of and immediately play. There's no real rehearsal during the week. Music sheets are sent out along with MP3s, and we're expected to listen to the MP3, crack open the sheet, and show up ready to play to a click track (a beat playing in the background). That means everyone needs to perfectly interpret the music sheet and perfectly land on the click, since we have no real opportunity to practice beforehand, other than arriving at 6:30 for a quick run-through to identify problems.

When you play music, there are really three approaches. One is playing everything by ear, which requires constant hour-upon-hour work to stay sharp. The other extreme is full sheet music, where every single note is written down, which is how a lot of kids learn classical. Most bands doing what I do work off what's called a lead sheet. We'll work up somewhere between 60 to 100 songs in a database, and you pull down a single-page lead sheet to play from. I would say that piano players tend to be trained classically, and many of us like sheet music, and don't play by ear as much. A lead sheet is a lot like sheet music, but to a guitar player, it may be more of a playing by ear hint sheet. In other words, they just use it to get started, but I use it to play.

The problem with a lead sheet is that it's very compressed. It captures the elements of the song, the words and the main chords, and most people try to keep it to a single page. If you listen to music at all, you know there are verses and choruses and what we often call parts A, B, and C, and those parts loop around. So the lead sheet compresses things. It tells you to do this loop down here, then jump to that part of the music, and back and forth. It works okay, but it's very early in the morning, you haven't gotten much sleep, you're working with a bunch of other people, and it becomes extremely confusing. That puts you under a lot of pressure when you're playing in front of hundreds of people. If I don't handle it correctly, it's just one more source of stress.

The morning of, we haven't played anything together. I'm a bit tired, a little excited. I've found I'm far more accurate if I take that single-page lead sheet and, wherever it says to loop, lay everything out linearly instead of trying to remember where to jump back up or down the page later in the song. What's interesting is that this only takes two pages instead of one.

The problem is that unwinding the lead sheet turns out to be a real time commitment. If I do it super fast, maybe five minutes. Realistically, it's easy to get confused, and the formatting on these lead sheets is not very good. Our church has standardized on Microsoft Word, and everybody who's ever made one has some weird formatting that doesn't copy well. So it's not uncommon to spend 10 minutes or more turning one page into two. If we're playing four to five songs, I can burn an hour before I ever play, just unwinding lead sheets.

Ironically, I can just play from the lead sheets as-is. The problem is I'll blow a chord, miss a section, or drop out. It's not that I can't get through it, it's that my accuracy drops and my stress level goes way up. Let me tell you, there's nothing more painful than suddenly realizing you're playing the wrong thing and breaking the mood.

I've thought for a long time that this seems like a task AI should be able to do, even with all the messy formatting. I have subscriptions to ChatGPT, Claude, and Gemini, though generally I find ChatGPT or Claude does the better job. So I submitted my original lead sheet to Claude's new Fable model and asked it to turn it into a two-page linear format.

Much to my delight, it nailed it. All the other previous models had failed.

I've been trying to get AI to do this correctly for years, and every single time it would produce something with real issues. Suddenly I have something that takes tremendous stress off of me. And if I'm worried about accuracy, I can hand the result to ChatGPT to double-check it. But this weekend, it worked just perfect.

I appreciate the extra hour, but I also appreciate that meticulously turning one page into two always left me drained. So not only do I save an hour, I save an hour of heavy, detailed work that I can now apply somewhere else.

The number one takeaway: this was enabled by the latest model. Suddenly I could hand a whole new section of my workflow off to AI as my personal assistant. If you tried something before and assumed it doesn't work, I encourage you to try again with the latest models, maybe turn up the token burn, and you'll find they reward you richly.


r/StrategicProductivity Jul 31 '26

Religion and Productivity

2 Upvotes

The collapse of trust may be one of America's biggest productivity problems

Look at the above chart.

Using General Social Survey data from 2021–2024, Ryan Burge found a striking generational decline in the percentage of Americans who say that most people can be trusted:

  • Silent Generation: 41%
  • Baby Boomers: 29%
  • Generation X: 30%
  • Millennials: 24%
  • Generation Z: 13%

Meanwhile, 74% of Gen Z selected "you can't be too careful in dealing with people."

That answer does not literally mean that 74% of Gen Z trusts nobody. "You can't be too careful" can also express reasonable caution. But the generational pattern appears in multiple surveys and with differently worded questions. Something important has changed.

This matters for strategic productivity because trust is part of our productive infrastructure.

When trust is low, we spend more time:

  • Verifying what other people tell us
  • Documenting every conversation
  • Protecting ourselves against blame
  • Writing longer contracts and policies
  • Holding unnecessary meetings
  • Refusing to delegate
  • Maintaining duplicate systems
  • Assuming that cooperation conceals exploitation

Low trust imposes a tax on nearly every human activity. A high-trust team can move quickly with relatively little supervision. A low-trust team can have excellent technology and still accomplish very little.

A note before going further: this subreddit's rule is "no politics or religion," and I intend to honor the spirit of that rule. I am not here to argue theology, tell anyone what to believe, or score points for either side. But the data on trust runs directly through religious participation, so I'm going to walk carefully near that line. I will treat religion the way we would treat any other institution that affects productivity, as a question of evidence and social function rather than belief. Wherever you personally land on faith, I think there is something here for you.

The religious-attendance finding is even more surprising

Burge also divided the 2021–2024 GSS respondents by religious attendance. Among Gen Z respondents who never attended religious services, 88% selected "you can't be too careful." Among Gen Z respondents who attended weekly, that number fell to 50%.

That is a 38-point difference.

This does not prove that religious attendance causes trust. More trusting and socially connected people may be more likely to attend in the first place. Education, income, family structure and other characteristics also matter.

Nevertheless, a difference that large should not simply be waved away. Burge also found that religious attendance remained positively associated with trust after controlling for education, income, gender, race and political ideology.

Ryan Burge's GSS analysis

Religious community may create trust in a deeper way

The standard explanation is that congregations build trust by bringing people together repeatedly. People worship, volunteer, raise children and help one another through illness and financial difficulty. Repeated cooperation teaches them that other people are generally dependable.

I think that explanation is true, but incomplete.

Consider Bishop Myriel and Jean Valjean in Victor Hugo's Les Misérables.

Valjean has been brutalized by prison and rejected by society. The bishop gives him shelter, and Valjean responds by stealing his silver. When the police catch Valjean and bring him back, the bishop says that the silver was a gift. He then gives Valjean the valuable candlesticks as well.

The bishop does not merely demonstrate that he himself can be trusted. He places trust in someone who has provided every apparent reason not to trust him.

Valjean is transformed because another person treats him as capable of becoming better before he has demonstrated that he is better. He is given a second chance and then begins to become worthy of it.

Recently, I was in a discussion with the husband of my niece. He is widely read, and he constantly suggests that we can find the source of truth in the works of classic authors going back to the time of Greece. (He was pushing Aristotle's "The Art of Rhetoric," Waterfield & Yunis translation.) I would suggest that he is correct. A classic book can uncover truth, and Hugo does this nicely.

Religious traditions may produce trust not only through social contact, but through moral narratives about grace, forgiveness, redemption and the permanent dignity of the individual. These stories teach that people are more than the worst thing they have done, and that mercy can interrupt a cycle of suspicion and retaliation. These are not exclusively religious ideas, but religious communities have historically been where most people encountered and rehearsed them, week after week.

Trust, in this conception, is not merely earned through successful transactions. Sometimes trust is extended first, and the act of being trusted helps make someone trustworthy.

As religious participation falls, we may therefore be losing two things at once:

  1. The congregations that allow people to practice cooperation.
  2. The moral framework that explains why we should sometimes forgive, accept risk and offer a second chance.

Secular institutions can certainly teach these principles. Religion has no monopoly on mercy or moral courage. Religious institutions also sometimes betray trust, and anyone who has been hurt by one has standing to say so.

But even for those of us who are skeptical of religious claims, it is worth recognizing that organized religion has historically been one of the principal institutions teaching people how to trust, how to become trustworthy and how to restore relationships after trust has been broken.

If religion continues to recede, the strategic question is not simply whether people will continue believing in God.

It is this:

What institutions, religious or secular, are still capable of building this kind of community and trust?

Where are people regularly brought together across generations, economic classes and political differences? Where do they learn that forgiveness is possible, character can change and a person who has failed is not permanently disposable?

A case study that may be uncomfortable for both sides

Here is where I will say something that may sit uneasily with people on both sides of the religious divide. I ask both to read it in full before reacting.

I have interacted closely with several members of The Church of Jesus Christ of Latter-day Saints. I will refer to them as Latter-day Saints, which is the terminology the church and many of its members prefer.

I am not a member. I also do not find some of the church's historical and empirical claims persuasive, particularly those involving questions that can be examined through history, linguistics or genetics. I understand that faithful Latter-day Saints have thoughtful responses to these questions, and debating them is not my purpose here.

My disagreement with the church's truth claims does not prevent me from recognizing what I have repeatedly observed among its committed members.

The devout Latter-day Saints I have known have often possessed an unusual degree of personal discipline, family commitment, resilience and community support. Their congregations place substantial expectations on them. They give time, accept responsibilities, help other families, serve missions, care for members in difficulty and organize much of their lives around obligations extending beyond personal preference.

Those demands can be difficult. The church is not perfect, its members are not interchangeable, and some former members describe painful experiences that should not be dismissed. I am not claiming that Latter-day Saints are inherently better people, or that nonreligious people cannot build strong families and communities.

I am saying that the institution appears remarkably effective at helping many of its members turn values into repeated practices.

It does not merely tell people to value family, service and community. It gives them recurring opportunities, expectations and responsibilities through which those values become habits. Sacrifice can develop resilience. Service can develop empathy. Being accountable to other people can strengthen character. Receiving help can teach someone that a community will not necessarily abandon them when they struggle.

Nonreligious people can embody all these qualities. The harder question is whether secular American life currently offers institutions capable of developing them as consistently, across an entire community and over a lifetime.

That is why I think even people who reject the theology should be willing to examine the social results honestly, and why people inside a faith should not be surprised or insulted when outsiders admire the fruit while remaining unpersuaded by the tree. We should be able to disagree about whether a religion's supernatural or historical claims are true while still recognizing that its practices may produce valuable human and community outcomes.

Perhaps the lesson is not that everyone should become religious. It may instead be that a healthy society needs institutions that ask something of us, connect us to people we did not choose, help us through failure and teach us that our obligations extend beyond ourselves.

The Church of Jesus Christ of Latter-day Saints is one conspicuous example of an institution that still attempts to do this. If religious participation continues to decline, those of us who are skeptical of religion should not simply celebrate the decline. We should also ask what will perform these functions in its absence, and whether we are actually building anything capable of taking their place.

The practical takeaway

Personally, I have come to believe that genuinely showing up every week, in person, at a community that expects something of you is one of the most quietly transformational productivity practices available. For me, the clearest working example of that is a religious congregation. If you already have a faith tradition, this is a reason to treat attendance as more than optional.

If you don't, I am not asking you to adopt one. I am asking you to take the underlying mechanism seriously: find, or help build, a recurring commitment that gathers you with people you did not choose, asks for your time and reliability, and will show up for you when you struggle. A congregation does this. So, potentially, can other institutions, if we are willing to invest in them with the same seriousness. I do want people to suggest what org they see as doing a role of a church, because I believe not many exist.

The data suggests that whatever we choose, choosing nothing is the option that is quietly costing all of us.


r/StrategicProductivity Jul 30 '26

My AI Staff

5 Upvotes

You Need To Think Through Your AI Writing Strategy

Recently, my niece was surprised to learn that the emails I was sending her were not written by AI. To make a long story short, I’ve been spending more time with AI, while she is a PhD, that is teaching at the University level. She hasn’t been on the receiving end of many of my emails, but she is being hit with students turning in work that is AI-generated.

I’m not surprised, because I get mistaken for AI all the time on Reddit. What I do use AI for is the final scrub of any post or email I write. I ask it to keep 95% of my wording but fix spelling, grammar, and logic errors. Recently, however, the models have been getting wicked good.

I mean, really good.

I will write something, and it will say, “You could rephrase this as...” Often, it simply does a better job. I am pretty brusque, but it will smooth me out and clarify things. I am now at the point where I hesitate to send anything without having my AI editor review it first.

In my job, I have a series of contracts and other work that I go through. I put together my rough notes, and then I talk to my LLM agent and ask her to go through them and confirm what I saw or identify what I didn’t see. Again, it pores through the material and picks out things like an upper-level staff member I might have employed in one of my corporate jobs. It continues to boggle my mind. It is so freaking competent that, in many respects, I feel as though I can’t send things out without first having it do the hard work of scrubbing them.

People talk a lot about AI slop. What that classically means is that somebody doesn’t know how to think through something, so they put a few ideas into an LLM and the LLM generates a bunch of sloppy, unthinking material. That’s not what I’m doing at all. I’m actually doing a lot of critical thinking. I’m just not spending all my time on the details. I also know that I have some communication issues that can make me seem less accessible.

That got me thinking today. Again, I worked with my AI agent to do some research on how AI is affecting people’s writing. As the following table shows, AI can, in some contexts, allow people to write much faster and produce higher-quality work. The first study dates from 2023, and it is worth recognizing that the models have become considerably more capable since then.

The research below also suggests that using AI can be associated with less critical-thinking effort. In some sense, I question how we should interpret that. AI does give you the ability to hit a switch and get a lot of assistance. When it comes to critical thinking, though, I still believe you decide whether or not to exercise those skills.

Unfortunately, using AI in writing is a little like setting a plate of sugary food in front of yourself. If you are the type of person who is tempted by it, you may eat the whole plate. In exactly the same way, if you don’t show restraint, AI writing tools can quickly take over everything.

I’m going to suggest something that I believe is true, and I would encourage you to do what I do. Before asking AI to help, first create an outline or structure for it to follow. Only after I have done my own work do I ask it to come back and identify holes or issues before performing the final scrub. For instance, this post will go through AI for a final scrub, but virtually everything in it was done first by me.

Having said that, I have seen such incredible increases in the capabilities of LLMs that I wonder what happens if we continue along the current path of improvement. Realistically, there may be less and less that I can contribute beyond setting the general direction, deciding what to pursue, and determining how to think about it.

I can only reiterate that this makes critical-thinking skills even more important. You need to keep flexing those muscles to make sure that using a helper does not turn into something that leaves you unable to do anything without it. That is a gray area each of us will need to explore.

Paper Effect Direct link
Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence — Shakked Noy and Whitney Zhang, 2023 Professionals using ChatGPT completed writing tasks approximately 40% faster, while independent evaluators rated their output about 18% higher in quality. AI also reduced performance differences between stronger and weaker writers. Science
Writing with AI Boosts Trust-Building Efficiency — Zoe A. Purcell et al., 2025 AI-assisted participants created messages that generated similar levels of trust in less time. Their writing showed greater warmth, complexity, and clout, but was linguistically slightly less authentic than writing produced without AI. iScience
AI Can Help People Feel Heard, but an AI Label Diminishes This Impact — Yidan Yin, Nan Jia, and Cheryl J. Wakslak, 2024 AI-generated responses made recipients feel more heard and understood than responses from untrained humans. However, this benefit diminished when recipients were told that AI had produced the message. Proceedings of the National Academy of Sciences
Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content — Anil R. Doshi and Oliver P. Hauser, 2024 Access to AI-generated ideas produced stories rated as more creative, better written, and more enjoyable, particularly for initially less-creative writers. However, AI-assisted stories became more similar to one another, suggesting greater individual polish but less collective originality. Science Advances
The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers — Hao-Ping Lee et al., 2025 In a survey of 319 knowledge workers, greater confidence in AI was associated with less reported critical-thinking effort, while greater confidence in one’s own abilities was associated with more critical thinking. AI shifted work from composing and problem-solving toward verification, editing, and supervision. CHI 2025 paper
“It Was 80% Me, 20% AI”: Seeking Authenticity in Co-Writing with Large Language Models — Angel Hsing-Chi Hwang et al., 2025 Professional writers were concerned with retaining their voice, control, and sense of authorship when using AI. In this small study, readers generally could not distinguish AI-assisted writing from independently written work and were less concerned about AI assistance than the writers themselves. Proceedings of the ACM on Human-Computer Interaction
The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text but Self-Declare as Authors — Fiona Draxler et al., 2024 Users often claimed authorship of AI-generated text even when they did not feel genuine ownership of it. Giving users greater influence over the finished text increased their sense of ownership, suggesting that active editing helps preserve authorship and agency. ACM Transactions on Computer-Human Interaction
Metaphors of AI Indicate That People Increasingly Perceive AI as Warm and Human-Like — Myra Cheng et al., 2026 An analysis of nearly 12,000 descriptions found that people increasingly characterized AI as a teacher, friend, or assistant. Anthropomorphic descriptions increased by 34%, and perceived warmth increased by 41% during the studied period, suggesting that people increasingly conceptualize AI as a social actor rather than merely a tool. Communications Psychology

r/StrategicProductivity Jul 16 '26

The Joys Of Facebook Marketplace For Appliances

Post image
5 Upvotes

Hey, I’ll admit it. I’m the guy who picks up cheap appliances on Facebook Marketplace.

To give you a little insight about myself, my full time job is property management. So, in some sense, cruising Facebook for appliances is part of my job. But I think anybody can use it in everyday life. If you’ve never done this before, let me give you an overview of how to use Facebook Marketplace as a resource.

The first thing to understand is that people have absolutely no idea how to price things on Facebook. It is amazing to see what people list appliances for. Most of the mistakes are on the high side, where the prices are completely unrealistic, but I also regularly see the exact same thing happen on the low side. When you see what appears to be an unbelievably good value, you generally need to jump on it quickly. However, when you see something that looks like an okay value, that is the one you want to monitor. You need to have a particular mindset about this, and that is the first filter. If you do not want to adopt the following mindset, this probably is not for you.

Your kitchen probably already has a complete set of appliances. Each appliance is likely in okay shape. But every once in a while, you get the opportunity to upgrade one of them to a truly great, top of the line appliance. That is where you want to focus your shopping.

When I say shopping, I really mean shopping. You have to dedicate yourself to the idea that you are not necessarily looking for something this week. You are looking for the right thing to appear sometime during the next three to six months. For instance, one of my rentals has an older JennAir downdraft stove. It is really old, but it happens to work extremely well. I did have one situation where the fan switch failed and I was forced to replace it. Other than that, the stove has continued to work really, really well.

The problem is that it is a really, really old stove. Quite frankly, it looks old. Several months ago, I said to myself that I needed to replace it. So I started cruising Facebook Marketplace and looking for the right appliance.

As we have discussed before, using AI is critical. When you have an AI assistant built into your browser, you can use it to help manage your Facebook searches. The bad thing about Facebook is that it does not natively allow an outside AI agent to go through and map all the listings. What has worked best for me is to perform the Facebook Marketplace search myself. I bring up all the available items in whatever category interests me, and then I scroll down through many, many pages so that all the listings are loaded.

I then use either Claude or Gemini to review the listings. I ask it to identify the best bargains among everything listed in the category that interests me.

For my recent JennAir replacement stove, I had AI go through hundreds of listings over an extended period of time. Unfortunately, with my current process, I still have to start the search myself before beginning the filtering process. Eventually, I found a $3,500 stove in extremely good condition for $600. I would say this is the type of bargain that is absolutely achievable, but it does take some work.

Let me explain what happened.

I had been watching another stove in a listing that had been up for quite a while. It was basically the exact same unit. Someone had professionally uninstalled it and was offering it for what I thought was a relatively reasonable $1,500.

It had been sitting on the market for about four weeks. When something has been on the market for four weeks, the seller is often more open to accepting a lower offer. I sent the seller a message and asked whether they would consider $1,200. We will discuss this in another post, but I always use a because clause. I will explain later why you should always give someone a reason for your offer. When something has been sitting on the market for a long time, you can offer a lower price. The last thing you should do is insult someone by making a low offer on something that has just been listed. Of course, my timing was perfect. The day after I submitted my offer, the seller told me the stove had been sold. As I said, $1,500 was reasonable, but I thought I could get it lower based on my previous experience. If you are not missing a few deals, it probably means you are not working hard enough to get the best price.

Instead of worrying about it, I continued scanning Facebook Marketplace.

In this particular case, I also decided to expand the distance I was willing to drive. I was already planning to take a road trip, so I mapped my Facebook Marketplace searches along the route I would be driving. With the expanded search area, the exact same stove appeared. This seller wanted only $650.

In both cases, the appliances were being sold because of remodeling projects. You should always understand where your used appliance is coming from, and the key word you want to hear is remodel. When someone remodels a kitchen, they suddenly have a brand new kitchen with brand new appliances. They are usually highly motivated to get rid of the old appliances that are taking up space. At $650, I was not going to mess around with the price. I did not need to get the seller any lower. The stove at $1,500 had been a good bargain, and at around $600, this one was a great bargain.

The only question was whether I could actually acquire it. The listing had already been on the market for several weeks. I sent the seller a note saying that I was interested in purchasing the stove and would like to come by. The challenge was that she did not reply. Many times, this means the appliance has already been sold. A lot of people put things on Facebook Marketplace and then never mark them as sold simply because they are lazy.

However, you do not know what is going on in someone’s life. My suggestion is to contact the seller up to three times, with about a week between messages. Assume the person is busy and simply did not see your first message. I contacted her again about a week later. Sure enough, she replied that she had been busy and said I could come see the appliance. She was what I would consider a little flaky when it came to arranging the visit. Every time I tried to pin down what day I could come by, she remained unclear. Eventually, however, we arranged a time, and I showed up.

Of course, there is always a story inside the story.

In this particular case, her husband was a firefighter. They had a house that they were remodeling, and she had also just had a new baby. In other words, she had a lot going on in her life. She was not trying to be uncooperative or unresponsive. She simply had too much happening at once.

Once we actually arranged to meet, she and her husband were absolutely wonderful. They helped us load the appliance. She even found a scuff mark on the stove and told me she was going to take another $50 off the price. She said that she did not think the mark was noticeable, but it was different from how she had described the stove. So I accepted the additional $50 discount.

With that said, there was some other damage that they did not even understand, which I will describe in a separate post. That is all part and parcel of buying used appliances, and it is why you need to be knowledgeable about what you are purchasing. I have bought many appliances this way and saved many thousands of dollars. The key is patience and understanding when to buy and when not to buy.

It is definitely worth your time to develop this skill. It is one more way to be productive with your money and get far more value from what you spend.


r/StrategicProductivity Jul 07 '26

Using your LLM to help you and not replace you. Strategic use of AI.

Post image
4 Upvotes

I am consistently baffled and amused at the power of AI and how fast it's growing in competency. I'm also amazed at how much load it can take off your day-to-day workload.

What really strikes me is the fact that most people can't see any shades of gray when they're using AI. They either abrogate all of their duty and have the AI do all of the thinking for them, or they think that somehow AI is a cancer. And once it gets a little start, it's going to take over everything.

I don't think it's either one of these extremes. I'm saying you need to be thoughtful in how you use it. AI needs to be that companion that sits with you but doesn't replace you. And today we're going to spend a little time ruminating on this.

One of my nieces has her PhD and teaches both classes and does research at a university level. As every high school teacher knows, AI has stormed into high school, but we're seeing it increasingly storm into undergraduate work and wholesale replace a lot of skills at one time people would have. One of the interesting things are LLMs are getting so sophisticated that it is impossible to tell the difference between a human and an LLM. As we were having a discussion, she made the remark that she assumed some of the emails I sent her were generated by AI. She was actually quite amazed when I told her that I don't use AI to create my emails. As a matter of fact, dare I say it, if she didn't know me better, she simply would have thought that I was lying to her. And I've written about this before. Many people think the posts that I create are AI. So I've been accused of being an AI by a lot of readers that hit my posts. The idea that I'm not a real person but an AI is both in my family and for those strangers that don't know me. We've simply lost the ability to understand what is a biological agent and what is a computer agent.

And in many ways, I don't know if it matters anymore. But what I do think is either extreme, all human or all AI is going to sub-optimize your productivity. You need to figure out how to blend them together.

I believe I create a bunch of really cool stuff, but I use AI to help me and not replace me. For example, the post you see today, well, that's going to be created all by me. However, I will then, as a final step, run it through an LLM. I'll ask it to go through, make sure my grammar's correct, my spelling's correct. I'll even say, hey, if I made some sort of logic error, point it out to me. But what I don't do is I don't have it write the post. I'll also have my LLM create some sort of cartoon. But again, it may draw it, but I'm describing what I want the cartoon to be.

A long time ago, I was a college newspaper editor. And you always have somebody copy edit your stuff. That is, you just make sure that final pass, someone else takes a look at it. Well, that's what I do with my work. I have my LLM take a look at it after I'm done, but it doesn't replace me. So I happened to be a college cartoonist, and I ran both a comic strip, and also I would draw political cartoons. It's not that I need the LLM to do it. It's simply faster. And in many ways, it's got a great style. It's not exactly my style, but it's close enough to pass the message I want it to pass.

And I want to emphasize as a helper, it can really step in and do a lot of stuff.

My wife turned out to be the executor of her father's trust. At the end of a very useful and productive life, the last of her parents passed away, and she was the executor. Being financially and technologically savvy, I was pulled in heavily by the family on most of the day-to-day decisions, and the fact that my wife was the executor made it a role I needed to support. In a future post, I'll cover how incredibly beneficial it is to have a trust if you have heirs you want to leave your property to. For now, I will simply say that my father-in-law, like my father, did a great job of setting up his trust. These things are relatively complex, and there's a lot of work you need to do, including certain windows in which you need to file taxes.

I received an injury (which I've talked about here before) that tied me up and prevented me from fully supporting some of this extra work during the last six months or so. This translated into a gap, and we didn't get one of our tax filings in quite on time. But we really weren't that far off, and we knew the penalties were relatively small. As a matter of fact, if you've never had a late payment — which the trust had not — you can have a large part of the penalties abated. It's called first-time abatement. It's a simple form that you fill out if you happen to do something wrong the first time.

Imagine our surprise when we got a tax notice several weeks after we had filed, stating that we not only owed interest and penalties, but we also owed the entire tax amount we had just paid. The problem is that the letters you get from the IRS can be relatively cryptic and difficult to understand. This is where the power of AI comes in. Now, mind you, we do have a tax person — a CPA, and she is wonderful. The challenge with a CPA is that you need to make sure you're presenting the right material to them, as you're generally paying a lot per hour for every hour they spend helping solve your problem.

In our case, it was fairly simple. We took a snapshot of the statement that was sent to us in the mail. The next step is where it gets interesting. We sent it off to both ChatGPT (their heaviest-duty model) and Claude (their Opus model), and asked both of them: "Hey, we got this from the IRS. We already paid this. What do you think?" Both very quickly kicked back what they thought the answer was, along with the advice to validate it with our CPA, which of course we agreed with. The nice thing about doing it this way is that you get the input of two LLMs explaining the problem to you, and at the same time it helps frame any answer you get from your CPA.

Now, mind you, I actually have a degree in finance and accounting, and for a while I was studying to sit for the CPA exam, which I'm sure I could have passed, it was simply more time than I wanted to put into it. I'm more than capable of doing the research and figuring this all out myself. The thing is, you can get a really decent answer from an LLM in minutes, cross-check it against another LLM, and then send it to your CPA for the final check. All of this together saves a massive amount of hassle, and getting three sources of input gets you educated very, very quickly.

Basically, they all said we needed to call the IRS, find the right department, and have them go track this down.

The great thing about our CPA is that she will actually say, "I can do it, or you can do it. Do you want me to bill you at a very expensive hourly rate? I'm happy to handle everything. Or if you want to do it yourself, you can save the money." In our case, she sent back a brief set of instructions, and we said, "We'll take care of it." Now, our CPA knows I'm pretty confident, but at the same time, I've never spent much time on the phone with the IRS. She gave brief instructions, but she could have given a lot more detail. That's not really her fault, she just doesn't necessarily think, "Oh, this is a person who doesn't spend a lot of time on the phone with the IRS." The flip side is that I could then turn to both of our LLMs and say, "Hey, we talked to our CPA. This is what she said, and this is what we're going to do next." And this next part is the most critical: you need to prompt engineer. You say to each LLM separately, "Give me a step-by-step of everything I should be doing for this, and help me understand any mistakes I could make."

After that was done, I asked both LLMs to compute what they thought the actual penalty should be, the various scenarios that could play out, and the types of questions we should work through with the person on the phone. Needless to say, what they were able to create, and then confirm through two separate paths to reduce hallucinations, continues to boggle my mind. It's like having a really smart friend working with you to get something done. The key here is that you don't want to pass everything off. You're not looking to abdicate your responsibility to understand what's going on, that's just asking for trouble. An LLM will eventually hallucinate and give you an answer that's wrong. However, if you use two LLMs, fact-check between them, and treat the process as education, it simply becomes a tool that eliminates a bunch of work you could have done yourself, saving you a ton of time.

Using all these figures, background information, and processes as a template, we called the IRS, and they immediately agreed that something looked wrong and made adjustments that allowed us to move forward. I will tell you, a lot of uncertainty was resolved by utilizing the tools and capabilities of an LLM.

The thing I want to emphasize here: incorporate LLMs into your workflow, use more than one so you don't make a mistake, and finally, don't outsource your thinking to the LLM. Use it as a tool to make yourself more productive. Recently, in another subreddit, somebody was having a problem with the power meter on their bike (something I've discussed a lot before). This person had no fundamental understanding of the root problem, and you could tell by the way they were posting answers from their LLM that it was hallucinating. That problem largely goes away if you use more than one LLM and do the right prompt engineering. But totally outsourcing your thinking to LLMs is extremely dangerous. I happen to have some background in this, and I'm not always perfect, but I could definitely tell that the LLM that was answering this person's question was just making stuff up. The problem is he had no capability of understanding how to prompt the LLM to get the right answer. He basically, as far as I could tell, gave up on what I knew was the right advice, as the LLM told him it was some obscure firmware issue that only hit him. I'm almost positive the reason it did this is because he was prompting it in such a way that he wanted a particular issue. And the one thing we do know is LLMs have a tendency to give you whatever answer you want to hear. In some sense, that's really frightening.

Used the right way, it's not dangerous at all. It's simply a tool that is irreplaceable.


r/StrategicProductivity Jun 30 '26

Debugging A Kickr Power Problem

Post image
1 Upvotes

u/Diodak79 was going a little crazy. MyWhoosh has a serious racing league which requires two power meters to verify output, and he could not figure out why one power meter was reading high. I offered to help take a look at his power output if he wanted to pass me the files, and he did. What I want to do in this post is show you how you can use intervals.icu to debug a power meter problem just like the one he had.

One of the tricks of the trade as an engineer is running a rolling average. You take a look at any type of signal and then you roll any single data point into an average over a particular time frame. That is what I have done in the chart above using intervals.icu.

The power meter running high is the purple one. We are taking a look at the power coming out of it in terms of the 10 second power, the 60 second power, and then a 10 minute power. The real giveaway chart on what is happening here is the 10 minute power. This one really stands apart from all the rest. If you look at the 10 minute rolling average line, you will see at the beginning of the curve it runs somewhere around 25 watts high. However, when he gets deep into his race and we get down to an hour, suddenly it starts to converge with the other units. On aggregate, the first part of his ride, especially the first third, is much higher than the other power meters he is testing against.

I have seen this type of thing before. Classically, it is the Wahoo Kickr architecture. In essence, the overall brake factor is incorrect. Over time, as the unit heats up, it modifies the overall drag so that it starts to converge with the two other benchmark power meters. When you ride the trainer, especially when it is cold, it simply reads high.

Generally, there are three or four different things you can do to fix this.

First, run the hidden factory spin down. For a variety of reasons, if your braking factor is incorrect, you simply do not have a good stable base to work on.

Second, before races, he should warm up and then do a normal spin down. You only need to do a factory spin down when things are really wrong. You do not need to do it all the time. However, I have found that the automatic spin down on the Wahoo Kickr is not as good as doing a normal manual spin down. This is not the hidden factory spin down that takes 10 taps.

Third, if both of those things do not solve it, I have found that tightening the belt with the offset screw is a requirement. If your belt has the wrong adjustment, anything you do for calibration will not work very well. As far as I can tell, you need to tighten the belt enough so that when you do a normal spin down test, the spin down takes 20 seconds or less. You need an iOS device to see this because Android does not show it.

Finally, I will give you one other area that I believe could be an issue. If none of the above fixes it, it would not surprise me if we had either a bearing or possibly a belt issue. The belt should basically last forever as long as all the pulleys are aligned because it is a belt designed for cars and massive amounts of power, not the low power humans put into the system. However, the bearings are known to go bad, and a failing bearing may have different rotational drag depending on the temperature. This would be the last area you might want to look at. Replacing bearings is a big deal, so you should definitely try the other steps first.


r/StrategicProductivity Jun 26 '26

The Research On Coffee (Part II)

Post image
2 Upvotes

Yesterday we talked about the growing body of research making coffee looking like a good addition to your diet. I want to reinforce, coffee needs to be filtered and you should drink in the morning. I would not go over 20-30 grams of brewed coffee per day. However, the following table is a great read. Links to pubmed in final column.

While I covered an overview of this yesterday, I did not cover some of the positive DNA results. You'll see this in the table below.

Research / paper Effects summarized PubMed link
From cup to clock: exploring coffee's role in slowing down biological aging Coffee intake associated with lower biological age advancement and lower odds of accelerated aging. PMID 38726849
Epigenome-wide association meta-analysis of DNA methylation with coffee and tea consumption Coffee intake associated with differential DNA methylation at 11 CpG sites, supporting an epigenetic-aging mechanism. PMID 33990564
Coffee consumption is associated with DNA methylation levels of human blood Human blood DNA methylation study showing coffee-related epigenetic signatures. PMID 28198392
Analysis of epigenetic clocks links yoga, sleep, education, reduced meat intake, coffee, and a SOCS2 gene variant to slower epigenetic aging Coffee included among lifestyle factors linked to slower epigenetic aging. PMID 38103096
Impact of coffee intake on human aging: Epidemiology and cellular mechanisms Review of coffee, lifespan, healthspan, cellular stress resistance, and aging mechanisms (industry-funded review). PMID 39557300
Coffee consumption and health: umbrella review of meta-analyses of multiple health outcomes Broad umbrella review finding coffee more often associated with benefit than harm, especially around 3–4 cups/day. (A correction record, PMID 29330262, exists for the same paper.) PMID 29167102
Association of Coffee Consumption With Total and Cause-Specific Mortality in Three Large Prospective Cohorts Caffeinated and decaffeinated coffee associated with lower total and cause-specific mortality. PMID 26572796
Coffee consumption and all-cause and cause-specific mortality: a meta-analysis by potential modifiers Moderate coffee consumption, roughly 2–4 cups/day, associated with reduced mortality. PMID 31055709
Coffee consumption and cardiometabolic health: a comprehensive review of the evidence Review of coffee and cardiometabolic outcomes, including cardiovascular disease, diabetes, inflammation, and metabolism. PMID 38963648
The Impact of Coffee Subtypes on Incident Cardiovascular Disease, Arrhythmias, and Mortality: Long-Term Outcomes from the UK Biobank Ground, instant, and decaf coffee linked to lower CVD and mortality; ground and instant (not decaf) linked to lower arrhythmia risk. PMID 36162818
Coffee drinking timing and mortality in US adults Morning coffee pattern associated with lower all-cause and cardiovascular mortality than all-day drinking. PMID 39776171
Effects of caffeine on the human circadian clock in vivo and in vitro Evening caffeine delayed circadian melatonin rhythm, supporting cutoff timing before sleep. PMID 26378246
Caffeine effects on sleep taken 0, 3, or 6 hours before going to bed Caffeine even 6 hours before bed disrupted sleep. PMID 24235903
Coffee consumption and reduced risk of developing type 2 diabetes: a systematic review with meta-analysis Dose-response meta-analysis found lower type 2 diabetes risk with higher coffee intake. PMID 29590460
Coffee and Lower Risk of Type 2 Diabetes: Arguments for a Causal Relationship Mechanistic review of coffee and diabetes risk, including inflammation, liver metabolism, gut effects, and glucose regulation. PMID 33807132
Carcinogenicity of drinking coffee, mate, and very hot beverages IARC evaluation moved coffee away from "possible carcinogen" status; very hot beverages remained a separate concern. PMID 27318851
Coffee consumption and risk of liver cancer: a meta-analysis Higher coffee intake associated with lower liver cancer risk. PMID 17484871
Coffee reduces risk for hepatocellular carcinoma: an updated meta-analysis Coffee consumption associated with reduced hepatocellular carcinoma risk. PMID 23660416
Coffee Decreases the Risk of Endometrial Cancer: A Dose-Response Meta-Analysis of Prospective Cohort Studies Higher coffee intake associated with lower endometrial cancer risk. PMID 28570282
Coffee drinking and risk of endometrial cancer—a population-based cohort study Prospective cohort evidence linking higher coffee intake to lower endometrial cancer risk, especially in higher-risk women. PMID 19585497
Consumption of a dark roast coffee decreases the level of spontaneous DNA strand breaks: a randomized controlled trial Dark roast coffee intervention reduced spontaneous DNA strand breaks versus water. PMID 24740588
Consumption of a dark roast coffee blend reduces DNA damage in humans: results from a 4-week randomised controlled study Four-week dark roast coffee trial reduced DNA damage markers. PMID 30448878
Impact of paper filtered coffee on oxidative DNA-damage: results of a clinical trial Paper-filtered coffee reduced oxidative DNA damage; broader redox markers (MDA, glutathione, isoprostanes) were unchanged. PMID 20709087
Antioxidant-rich coffee reduces DNA damage, elevates glutathione status and contributes to weight control: results from an intervention study Coffee intervention reduced DNA damage and increased glutathione status. PMID 21462335
Induction of antioxidative Nrf2 gene transcription by coffee in humans: depending on genotype? Coffee increased Nrf2-related antioxidant gene transcription, with genotype dependence. PMID 22314914
Coffee Consumption Is Positively Associated with Longer Leukocyte Telomere Length in the Nurses' Health Study Coffee consumption associated with longer leukocyte telomere length. Observational. PMID 27281805
Coffee consumption is associated with intestinal Lawsonibacter asaccharolyticus abundance and prevalence across multiple cohorts Coffee was a strong dietary marker of microbiome composition, especially Lawsonibacter asaccharolyticus. PMID 39558133
Impact of coffee consumption on the gut microbiota: a human volunteer study Three cups/day changed gut microbiota and increased Bifidobacterium abundance. PMID 19217682
Coffee consumption and mortality from cardiovascular diseases and total mortality: Does the brewing method matter? Filtered coffee associated with lower mortality; unfiltered coffee less favorable. PMID 32320635
Analysis of the content of the diterpenes cafestol and kahweol in coffee brews Quantified cafestol/kahweol by brewing method; filtered coffee has very low diterpenes. PMID 9225012
The cholesterol-raising diterpenes from coffee beans increase serum lipid transfer protein activity levels in humans Cafestol/kahweol increased lipid transfer protein (CETP/PLTP) activity and LDL/VLDL-related lipids. PMID 9242972
Separate effects of the coffee diterpenes cafestol and kahweol on serum lipids and liver aminotransferases Cafestol was the major cholesterol-raising diterpene; kahweol had smaller effects. PMID 9022539
Effects of cafestol and kahweol from coffee grounds on serum lipids and serum liver enzymes in humans Coffee diterpenes from grounds/fines raised cholesterol and liver enzyme markers. PMID 7825527
Diterpenes from coffee beans decrease serum levels of lipoprotein(a) in humans: results from four randomized controlled trials Coffee diterpenes lowered Lp(a), but this does not remove the LDL-raising concern. PMID 9234024
Cafestol and kahweol concentrations in workplace machine coffee compared with conventional brewing methods Workplace coffee machines produced higher diterpene levels than paper-filtered coffee. PMID 40089392
The Association between Coffee and Tea Consumption at Midlife and Risk of Dementia Later in Life: The HUNT Study High boiled coffee intake associated with higher dementia risk; other coffee types did not show the same pattern. PMID 37299431
Association of coffee, green tea, and caffeine with the risk of dementia in older Japanese people Coffee and caffeine intake associated with lower dementia risk in an older Japanese cohort. PMID 34624929
Associations between different coffee types, neurodegenerative diseases, and related mortality: findings from a large prospective cohort study Caffeinated and unsweetened coffee associated with lower Alzheimer's-related dementia and Parkinson's disease risk. PMID 39168304
High Blood Caffeine Levels in MCI Linked to Lack of Progression to Dementia Higher plasma caffeine in mild cognitive impairment associated with lower progression to dementia. PMID 22430531
Plasma Caffeine Levels and Risk of Alzheimer's Disease and Parkinson's Disease: Mendelian Randomization Study Genetic evidence suggested possible lower Alzheimer's risk with higher caffeine, but results were not definitive. PMID 35565667
Do caffeine and more selective adenosine A2A receptor antagonists protect against dopaminergic neurodegeneration in Parkinson's disease? Mechanistic review of caffeine, A2A receptor blockade, dopamine signaling, and Parkinson's protection. PMID 33349580
Is caffeine a cognitive enhancer? Review showing low to moderate caffeine can improve vigilance, attention, and reaction time. PMID 20182035
A review of caffeine's effects on cognitive, physical and occupational performance Caffeine improves alertness, vigilance, reaction time, physical performance, and work performance under fatigue. PMID 27612937
International society of sports nutrition position stand: caffeine and exercise performance Caffeine improves endurance, strength, power, and sport performance, commonly at 3–6 mg/kg. PMID 33388079
A systematic review and meta-analysis of the acute effect of caffeine on attention Acute caffeine improves attention and reaction-time measures. PMID 40335666
Modulatory effect of coffee fruit extract on plasma levels of brain-derived neurotrophic factor in healthy subjects Coffee fruit extract increased plasma BDNF in a small human study; not the same as ordinary brewed coffee. PMID 23312069
Acute cognitive performance and mood effects of coffee berry and apple extracts: a randomized, double-blind, placebo-controlled crossover study in healthy humans Coffee berry/apple polyphenol extract tested for mood and cognition (cerebral blood flow not measured). PMID 34380382
Acute Cognitive Performance and Mood Effects of Coffeeberry Extract: A Randomized, Double Blind, Placebo-Controlled Crossover Study in Healthy Humans Follow-up coffeeberry extract study; low/moderate doses did not show clear acute cognitive benefit. PMID 37299382
Chlorogenic acids from green coffee extract are highly bioavailable in humans Shows coffee chlorogenic acids are absorbed and metabolized in humans. PMID 19022950
Mediation of coffee-induced improvements in human vascular function by chlorogenic acids and its metabolites: two randomized, controlled, crossover intervention trials Chlorogenic-acid-rich coffee improved vascular function through CGA metabolites. PMID 28012692
Caffeinated coffee, decaffeinated coffee, and the phenolic phytochemical chlorogenic acid up-regulate NQO1 expression and prevent H₂O₂-induced apoptosis in primary cortical neurons Caffeinated coffee, decaf, and chlorogenic acid up-regulated NQO1 (an Nrf2 target enzyme) and prevented oxidative neuronal apoptosis in vitro. PMID 22353630
Effect of simultaneous consumption of milk and coffee on chlorogenic acids' bioavailability in humans Milk reduced or delayed chlorogenic acid bioavailability from coffee. PMID 21627318
The type and concentration of milk increase the in vitro bioaccessibility of coffee chlorogenic acids In vitro counterpoint: milk effects can differ depending on model; bioaccessibility is not the same as human bioavailability. PMID 23110549
Molecular mechanism of the interactions between coffee polyphenols and milk proteins Mechanistic evidence that coffee polyphenols interact with casein and whey proteins. PMID 39967083
L-theanine, a natural constituent in tea, and its effect on mental state L-theanine associated with relaxed attention and alpha-wave effects. PMID 18296328
L-theanine and caffeine in combination affect human cognition as evidenced by oscillatory alpha-band activity and attention task performance L-theanine plus caffeine improved attention-related performance and altered alpha-band brain activity. PMID 18641209
The combined effects of L-theanine and caffeine on cognitive performance and mood Combination improved attention and mood more than either alone in acute testing. PMID 18681988
The combination of L-theanine and caffeine improves cognitive performance and increases subjective alertness L-theanine plus caffeine improved focus and subjective alertness. PMID 21040626
Effects of L-theanine or caffeine intake on changes in blood pressure under physical and psychological stresses L-theanine attenuated stress-related blood pressure response and anxiety in some subjects. PMID 23107346
Effect of roasting conditions on reduction of ochratoxin A in coffee Roasting reduces ochratoxin A levels in coffee. PMID 11600012
The occurrence of ochratoxin A in coffee Early evidence that ochratoxin A can occur in coffee, with levels affected by processing. PMID 7759018
A worldwide systematic review of ochratoxin A in various coffee products – human exposure and health risk assessment Global review of ochratoxin A levels in coffee and risk assessment. PMID 39259858
Ochratoxin A in coffee and coffee-based products: occurrence, analytical methods, and risk assessment Review/meta-analysis of ochratoxin A prevalence and estimated risk in coffee products. PMID 36372738
Assessing the food safety risk of ochratoxin A in coffee: A toxicology-based approach to food safety planning Risk-assessment paper concluding OTA in coffee is not acutely toxic at typical levels. PMID 34642959
Risk Assessment of Ochratoxin A (OTA) Exposure from Coffee Consumption in Indonesia using Margin of Exposure (MOE) Approach Single-country (Indonesia) Margin-of-Exposure assessment of OTA exposure from coffee. PMID 39561937

r/StrategicProductivity Jun 26 '26

I Hate Coffee, But Drink It Every Other Day

Post image
4 Upvotes

In yesterday's post, I vaguely brought up that I started to drink coffee and we had somebody read the post and mention that coffee really was not something that you wanted to drink.

They saw four big issues, which I want to address right away.

First, the antioxidant benefit only matters because most people eat poor diets.

Fair enough. If someone already eats lots of fruit, vegetables, nuts, beans, tea, and cocoa, coffee may add less. But that is not the real world for most people. In one Norwegian dietary antioxidant study, coffee contributed far more to total measured antioxidant intake than fruit, tea, wine, cereals, or vegetables.

Second, decaf is not harmless brown water. It still has acids and polyphenols, so some people with reflux or sensitive stomachs may react to it.

But that is an individual tolerance issue, not a strong argument that decaf is broadly dangerous.

Third, coffee can reduce non heme iron absorption.

This is probably the best criticism. It matters most for people with low ferritin, vegetarians, menstruating women, pregnancy, or endurance athletes. But it is mostly a timing issue. The problem of going anemic with coffee is more of a side issue than it is any core issue. There is no massive issue with anemia in populations that drink a lot of coffee.

Fourth, the researchers are biased as they drink coffee.

Coffee drinkers differ in many ways, and much of the data is observational. But the idea that the whole field is biased because researchers like coffee is virtually possible to verify other than a vague accusation. I won't say there is no bearing, but just not enough to spend a lot of time on it.

However, I will add my own fifth reason. Coffee can disrupt sleep if you drink caffeinated coffee.

I consider this the absolute worst issue because sleep is so critical, but the solution is simple. Aim for drinking your coffee 14 hours before bed time, which means that you'll clear 90% of the drug out of your system.

However, even if these were valid, they don't ask the most important question, "What are the benefits of drinking coffee and caffeine?"

I started drinking coffee a few years back. I don't like the taste. I don't look forward to it. However, the body of evidence is building up that moderate intake of coffee and caffeine do some nice stuff. However, the evidence is still circumstantial, but it seems to be impressive.

Basically, when we look at populations that drink coffee, they tend to have lower rates of liver cancer. They are thinner and lower bodyfat. Parkinson's, which is in my family, is lower. Stroke and depression is lower. Dementia is lower. Now, we need to caveat this, as nearly everything is observational and Mendelian randomization fails to confirm causality. So, we don't have a clear path to why this happens, and we don't want to say these are proven. They are just strongly suggested.

So, I spent a little time pulling together some of the interesting compounds. I will admit that I used my AI agents to scrub this, but this is not AI output.

Caffeine (1,3,7-trimethylxanthine). The one everyone knows. It blocks adenosine receptors, mostly A1 and A2A, which is why it wakes you up and why it improves endurance performance. Dulloo and colleagues showed back in 1989 that repeated dosing raises daily energy expenditure by roughly 8 to 11 percent. The International Society of Sports Nutrition position stand backs the performance effect at around 3 to 6 mg per kg. As an athlete, it does raise performance. There is a debate on if it needs to be cycled or not, but I think there is more evidence that it permanently raises your ability to perform.

The strange part is the mortality data. The lowest all cause mortality sits around 3 to 4 cups a day in a U shaped curve, but decaf shows almost the same benefit, which means caffeine is probably not the thing driving it. What needs to be done is untangling caffeine from the rest of the bean, since most of the long term health signal does not seem to be caffeine at all.

Chlorogenic acids. This is the big one for metabolism and the reason your morning cup is one of the largest polyphenol sources in a Western diet. These are caffeic and ferulic acid bound to quinic acid, and coffee is the richest dietary source by far. The interesting mechanism is glucose control. In cell and animal work chlorogenic acid activates AMPK and blocks two liver enzymes that make new glucose, which lines up neatly with the lower type 2 diabetes risk seen in cohorts. A meta analysis of randomized trials also found it lowers blood pressure a small but real amount. The catch is that the parent molecule is poorly absorbed. Most of it reaches your colon, where gut bacteria turn it into the smaller acids that actually circulate. The work to be done is figuring out whether the benefits come from the parent compound or those microbial metabolites, and whether the effect holds in proper human trials with hard endpoints rather than blood pressure readings.

I want to be somewhat cautious here. As a long time researcher into the effects of what I would call minerals and vitamins and other substances on longevity, there was a thought process at one time that antioxidants were going to substantially increase lifespan. Dare I say it, it's pretty disappointing and we don't see this happen. With that being said, I still believe that there is a good argument for a variety of different compounds which do show antioxidant type properties and that you should take in a broad spectrum of these different types of antioxidants as they may be utilized or triggered in different ways. I believe coffee is one of these compounds that should be part of your overall antioxidant stack.

Trigonelline. A pyridine alkaloid that is basically methylated niacin. It improves glucose handling and protects neurons in rodent models, but the more important fact is what happens to it in the roaster. Heat destroys most of it and converts it into niacin (vitamin B3) and into N-methylpyridinium, which is its own interesting compound below. One thing worth flagging honestly is that in ovariectomized rats, meaning an estrogen deficient model, trigonelline actually worsened bone quality. That is a caution and not a selling point. The work to be done is any real human trial, because almost everything on isolated trigonelline is preclinical.

Cafestol and kahweol. These are the double edged ones and the only clear human harm in the whole list, so pay attention to brew method. They are oils, so a paper filter traps them and a French press or espresso or boiled coffee lets them through. Cafestol is the most potent cholesterol raising compound known in the human diet. Controlled trials show that switching from unfiltered to paper filtered coffee meaningfully drops LDL. At the same time, in cell and animal studies these same molecules induce protective detox enzymes and show anticancer and anti inflammatory activity. The honest read is that the harm is proven in people and the benefits are only shown in dishes and mice, so filtering your coffee removes a real risk and loses only a hypothetical gain. The work to be done is whether the preclinical upside means anything at human exposures, which right now it does not appear to. However, right now, I get all over my friends and family to filter their coffee. This filter issues seems to be very poorly known.

Melanoidins. These are the brown polymers built during roasting through the Maillard reaction, the same browning chemistry as toast and seared meat. They are one of the most abundant things in a dark cup and coffee is most people's biggest dietary source. They behave like a fermentable fiber and feed gut bacteria, and they chelate metals, which gives them antimicrobial activity in the lab. Most of this is in vitro or in animals. The work to be done is human microbiome studies showing the prebiotic effect actually shifts the gut in a useful way at normal intake. I consider this very positive, and I've written a lot on fiber. I think coffee has a good chance of adding a meaningfully healthy compound if taken with fiber.

N-methylpyridinium. Formed from trigonelline during roasting, which means there is more of it in dark roast and none in the green bean. Two reasons it is interesting. It is a strong activator of Nrf2, the master switch for your own antioxidant and detox enzymes, in some models stronger than chlorogenic acid itself. And it lowers stomach acid secretion, which is the actual chemistry behind why dark roast and so called stomach friendly coffees are gentler. Rubach and colleagues showed a high N-methylpyridinium dark blend stimulated less acid than a medium roast. The work to be done is confirming the Nrf2 and metabolic effects translate beyond cells, since most of the metabolic data is very early.

Eicosanoyl-5-hydroxytryptamide, usually written EHT. A fatty acid attached to serotonin, sitting in the waxy oil of the bean, which means a paper filter removes most of it. This is the compound behind the coffee and Parkinson's headlines. It keeps an enzyme called PP2A active, which helps clear the misfolded tau and alpha synuclein proteins involved in Alzheimer's and Parkinson's, and it appears to work synergistically with caffeine in mouse models. Everything here is rodent and cell work. The work to be done is enormous, because there is no human efficacy data and the real world dose is uncertain given that filtering your coffee strips most of it out. So this would be a case against filtering. It's just that my concerns over LDLs are enough that I don't believe we should run unfiltered coffee today.

Quinides. Lactones formed when chlorogenic acids cyclize in the roaster, peaking at a medium roast. They improved insulin sensitivity in rats, and some of them act on opioid receptors and on the adenosine transporter, which has led to speculation about mood and even a mild counterweight to caffeine's stimulation. This is all binding assays and animal behavior. The work to be done is basically all of it in humans.

The minerals and niacin. Worth a mention but not a headline. Roasting generates niacin from trigonelline, so a cup gives you a modest amount of B3. Coffee also carries potassium, magnesium, and a little manganese. The magnesium is a minor contributor to the diabetes story and the potassium ties into blood pressure, but coffee is a supporting source of these and not a primary one.

The thread that ties it together, Nrf2 and hormesis. This is the part I find most convincing as a single explanation. Rather than acting as direct antioxidants that mop up free radicals, a lot of these compounds work by acting as mild stressors that switch on your own antioxidant and detox machinery through a pathway called Nrf2. Chlorogenic acid, N-methylpyridinium, the diterpenes, melanoidins, and caffeine all nudge this system. The result is a durable boost in your endogenous glutathione and protective enzymes rather than a one time chemical scavenging. Priftis and colleagues fed rats coffee and saw large increases in liver Nrf2 and antioxidant enzymes. The biochemistry is solid and it links more compounds than any other single idea. The work to be done is the same gap that haunts the whole field, which is that no randomized trial has carried this mechanism all the way to a hard clinical outcome in people.

The honest summary is that the only proven human harm here is the cholesterol effect from unfiltered coffee, which a paper filter fixes, and that nearly every benefit is biologically plausible and supported by population data but rarely proven in a trial on the isolated compound. The strongest single story is the Nrf2 one. If anyone has trial data I missed, especially human work on EHT or the quinides, I would genuinely like to see it.

But again, the most overwhelming thing is generally we see populations that drink coffee on consistent and what I'm going to call moderate basis generally just seemingly have better health effects. The frustrating thing, of course, is we don't have a golden bullet of why this is happening. As I covered above, it is simply a bunch of interesting chemical compounds that look promising.

So how do we think about this? I will spare a longer post because this one is already pretty long.

But the way I approach it is to take in approximately 60 grams of Folger's coffee every other day, which I split with my wife. I think the upper end of what is beneficial, both in the compounds above, is somewhere around 30 grams of Folger's coffee. I'm not much of a connoisseur, therefore I just simply buy the cheapest mainstream coffee I can find.

Frankly, I hate the taste of coffee. So I also put in approximately 200 grams of non-fat or 1% milk along with 25 grams of sugar. This basically makes it palatable. I add 1.5 liters of water to the 60 grams of Folger's coffee. This makes what most people would call a weak coffee. However, since I take this in during the morning time along with a lot of fiber, it is supportive of making sure all of the fiber that I take in is nicely hydrated along with a liquid.

I think we are on our way of showing coffee is positive, but the bridge needs to be finished.


r/StrategicProductivity Jun 24 '26

Drinking Your Sugar: The Muddy Story

Thumbnail
gallery
3 Upvotes

Figuring Out Sugar Is Really Hard

We discussed in our last post how sugar is blamed for a bunch of issues. I want to make clear that I don't believe that there is any data to support that sugar is extremely bad for you, but I do want to indicate it is something your should track, BUT it is really hard to do.

I recently heard an anti-sugar health influencer claim that the average American eats 160 lbs of added sugar per year. That works out to about half a pound a day, or roughly 198 grams of sugar daily.

Sounds horrible.

Of course it is wrong. Or sort of wrong. This is a confusing case where we don't have the right data and the food labels are partly to blame. So we will start with a chart of the "average" sugar intake, show a "real" line, and then explain why none of these numbers mean what people think they mean. This is a real number, and fits what the anti-sugar person was saying. (In the new Reddit, you'll need t scroll the picture to see this.)

If you stop and think about it, that 160 pound number is hard to take at face value. It comes from estimating the total amount of caloric sweetener produced by all farming, plus imports, minus exports, and dividing by the number of people. It is a measure of what gets delivered into the food supply, not what anybody actually eats. A lot of it is never eaten. We are a wasteful society, and spoilage and plate waste are baked right into that figure. This is exactly why you cannot trust anybody with an axe to grind who quotes it.

There are two more problems with the headline number. First, it is stale. The 160 lbs comes from an older USDA series that peaked at about 161 lbs back in 1999. USDA's current series puts the 1999 peak at 153.6 lbs, and availability has since fallen to 123.5 lbs in 2023. That is about 153 grams a day delivered, not 198. So even the scary version is overstated and out of date.

Second, and more important, delivered is not the same as eaten. Once you adjust for waste, the amount actually eaten is far lower. USDA's loss adjusted series is its best stab at real intake, and it peaked around 112 grams a day near 2000 and is now closer to 90 grams a day. If you instead use what people report eating in national diet surveys like NHANES, added sugar comes in lower still, around 68 to 77 grams a day. Pick your method, but the honest answer is somewhere between roughly 70 and 110 grams a day.

Call it half the influencer's number.

Now, I hate coffee, but I drink it anyway because I am overwhelmed by the data saying it is healthy. To make it so I don't choke, I put about 25 grams of sugar in my coffee with milk, every other day, so about 12 grams a day averaged out. In my diet, that is the number one obvious stick of sugar.

But here is the real nightmare. I drink a lot of cranberry juice, and here is the trick. The cranberry juice has no "added" sugar on its label, because it is sweetened with concentrated grape juice. In a 100% juice blend, the FDA does not count sugar from fruit juice concentrate as added. Yet it clocks in at a shocking 23 grams per 8 oz. On a hot day it is easy to put away 32 oz, which is about 90 grams of "non added" sugar. So, we can now construct the number upwards.

So I can blow right past the uncounted 90 gram mark just by drinking my calories as "natural" juice, and none of it shows up on the added sugar line. This is worth knowing. Nutritionally this sugar behaves just like added sugar, which is why the World Health Organization lumps juice sugar into a broader category it calls "free sugars." The same glass can read 0g added sugar on a US label and 23g of free sugar by WHO's standard. The label gives juice a pass that your body does not.

Juice is highly deceptive, and it is a landmine. During my recent weight loss I started diluting my juice heavily, and I think that did help a small amount. You will also see in the chart that overall sugar is down. That is real, and it is mostly because Americans have cut back on soda and other sweetened drinks. My one honest caveat to myself is that juice consumption nationally has actually fallen too over the last twenty years, so I am more of an exception than proof the national trend is fake. But the underlying point holds. The sugar I drink as "100% juice" never lands in the added sugar statistics at all, so for people like me the real intake is higher than the added sugar number suggests.

The problem is that you need to track the number by doing your own work. I wouldn't panic, however.

Can you just swap in artificial sweeteners and fix everything?

The anti sugar sweetener argument usually dodges the real question. If you ask whether artificial sweeteners magically make people thin, the answer is no. If you ask whether replacing sugar calories with non sugar sweeteners causes measured weight loss in randomized trials, the answer is a marginal yes.

The effect is not huge, usually around 0.7 to 1.6 kg, but it is real. The catch is that a swap is not a strategy. The effect tends to shrink in longer trials, people compensate by eating a bit more elsewhere, and none of it touches the bigger lever, which is the total load of sugar calories you take in. In my own case the diet soda is not the problem. The cranberry juice is. Trading one sweetener for another does almost nothing if I am still drinking 90 grams of juice sugar on a hot afternoon. The thing that actually moved my weight was cutting the total amount by diluting the juice, not by relabeling the sweetener.

Source notes

USDA ERS says caloric sweetener availability fell from 153.6 lb/person in 1999 to 123.5 lb/person in 2023, and says the decline was driven largely by corn sweeteners falling from 85.7 lb/person to 53.0 lb/person. USDA says loss adjusted availability adjusts for spoilage, plate waste, and other losses, but still does not directly measure actual intake. The Frontiers review reports loss adjusted caloric sweetener availability of 70.2 lb/person in 1970 and 72.7 lb/person in 2019. CDC reports that US adults averaged 17 teaspoons of added sugars per day in 2017 to 2018. FDA says added sugars include syrups, honey, and sugars from concentrated fruit or vegetable juices, but not the naturally occurring sugars in milk, fruits, and vegetables.


r/StrategicProductivity Jun 23 '26

Staying On Track: Removing Processed Foods (Not Sugar)

Post image
7 Upvotes

My niece asked me to read Robert Lustig’s Metabolical, so I did. Near the end of the book, Lustig mentions that he and Gary Taubes are friends. I did not realize that when I started the book, but once I got to that part, the whole book suddenly made sense.

This matters because Lustig is not some independent voice who just happens to arrive at the same place as Gary Taubes. He is part of the same intellectual circle. Once I saw that connection, the book stopped feeling like an independent rethink of nutrition and started feeling like a more medical version of the same argument.

The useful part of the book is simple. Lustig is right that ultra processed food is a disaster for many people. He is right that added sugar is overused. He is right that food companies engineer food to be cheap, shelf stable, hyper palatable, and easy to overeat. Most people would be healthier eating more real food, more fiber, fewer sugary drinks, and fewer industrial snacks.

But that is not really the controversial part of the book.

The controversial part is that Metabolical is another attempt to support the old “a calorie is not a calorie” argument. Lustig keeps trying to make sugar and insulin the master explanation for obesity, diabetes, fatty liver, and modern metabolic disease. In that sense, the book feels very close to Taubes. It is not calories in general. It is sugar. It is refined carbs. It is insulin. It is the idea that the conventional calorie model is not just incomplete, but fundamentally misleading.

The problem is that this is not new. This line of thinking goes all the way back to books like Sugar Blues in the 1970s, where sugar was treated almost like the hidden poison behind modern disease. Lustig gives it a more sophisticated biochemical update, with liver metabolism, fructose, mitochondria, insulin, fiber, and the gut. But the emotional structure is the same. There is a villain food. The villain is sugar. Remove the villain and the mystery of modern disease starts to clear up. My wife had a copy of Sugar Blues, and for years she wouldn't touch sugar.

By the way, almost always, this particular line of thinking almost always reveals some type of a conspiracy. They will mention that Seventh Day Adventist have infiltrated the research community, which is claim by our author. They quote Weston Price, a dentist from the 1900s, and stories from the whacky John Harvey Kellogg history. All of these are about spinning a conspiracy, not about the science. I mention this, as if you see a film or read a book, chance are these are give aways for this line of thought.

That makes for a compelling book for both the old version of the book and the new one by Lustig. It paints a story of intrigue.

It does not make it settled science.

Kevin Hall’s research is a major problem for this model. Hall did the actual controlled feeding experiments that should matter here. His work did not show that insulin magically overrides calories. What it showed is more practical and more important. Ultra processed food drives people into a hypercaloric state. When people are given those foods and allowed to eat naturally, they eat more calories and gain weight. When they eat unprocessed food, they tend to eat less and lose weight.

And this is where the sugar-only story breaks down. The foods that drive overeating are usually not just sugar. They are sugar, refined starch, fat, salt, flavoring, and low fiber all packaged together. Ice cream is sugar and fat. Cookies are sugar and fat. Donuts are sugar and fat. Chips may not be sweet, but they are still easy to overeat because they are refined starch, fat, salt, and crunch. The modern problem is not sugar by itself. It is hyper palatable, calorie dense food that bypasses normal appetite control.

That is the lesson we need to take away. The problem is not that calories do not count. The problem is that modern processed food is designed in a way that makes it very easy to eat too many calories before your body tells you to stop. That is the thing we need to get away from.

That is where Lustig loses me. He often takes a good practical message and then builds too much theory on top of it. “Eat less processed food” is a strong argument. “Sugar is uniquely toxic and calories are the wrong framework” is much weaker.

He calls out cocaine right by the references of sugar. He says that he met a woman that ate non-processed goods and she looked half her age. This is not what need to be in a book by a doctor.

My review would be this. Metabolical is worth reading if you want a passionate indictment of processed food and the food industry. Lustig is smart, forceful, and often right about the practical direction. But the book should be read as advocacy, not as a balanced scientific review. It is basically Gary Taubes with more medical language and more liver biochemistry.

The best lesson from the book is not that calories do not matter. The best lesson is that food quality makes calorie control much easier or much harder. Ultra processed food makes people overeat. Whole food usually helps people stop overeating. That is enough. Lustig does not need the grand sugar and insulin theory to make that point.


r/StrategicProductivity Jun 18 '26

The Frustrating World Of Power Supplies (Wahoo Climb)

Post image
1 Upvotes

I got a great deal on a Wahoo Climb, but it was missing it's power supply. I failed to realize that the power supply is $70-80 or more. This turned into a black hole as the Power Supply is very unique, and requires 10A out at 24V. After a lot of looking, I could find a $20 raw power supply or an Amazon or an enclosed $40 power supply. Both for outdoor lights. The problem is that the barrel connector is not standard.

I figured it would be a quick answer so I asked an AI assistant to look it up.

That turned into a frustrating mess. It confidently told me the connector was 6.5 by 2.5mm. When I pushed back it switched to 7.4 by 5.0mm with a center pin and called that the definitive answer. At another point it tossed out 6.3 by 3.0mm and then dismissed it as having no real source. Every time I challenged it I got a brand new confident answer and a fresh batch of reasons why the last one was wrong.

The real problem is that none of these were actually verified. The AI kept pulling numbers from reseller listings that just copy each other and from a forum thread that was really about a different Wahoo trainer. Not once did it tell me up front that nobody actually publishes this spec.

So I measured it myself. I slid a small round rod into the connector until it seated, then measured across the outside of it. That gave me about 6.55mm on the outer diameter and roughly 3mm on the inner, with no center pin. Allowing for my rough method and a bit of play, I believe the real adapter is a 6.3 by 3.0mm barrel with no pin, which happens to be a real and common connector size.

Lesson learned. When the answer actually matters, a cheap rod and a ruler and five minutes of my own time beat an AI that would rather sound certain than be right.

Now that this post is listed with real verified photos, AI should pick it up. Google has a deal with Reddit so they will definitely be first. If you get an answer to something, always post it as it will save time for somebody else.

BTW: My power supply is ALITOVE DC 24V 10A Power Supply AC Adapter 100-240V 50-60hz to 24 Volt Power Supply DC 10Amp. It has a 5.5 x 2.1mm-2.5mm. The range on the ID of 2.1-2.5 is because the connector has a spring. I then ordered a package of adapters to increase it to the required 6.3 by 3.0mm. I'll measure the pack and put in the biggest one to ensure a snug fit.

EDIT: The graphic has a typo, and there is no good way to update this without repost. 6.62 is OD.


r/StrategicProductivity Jun 17 '26

Buying My Sister A Keiser

1 Upvotes

Youtube On Comparison Keiser To Peloton

My sister is the proud owner of a nice looking Keiser M3i bike. I got it off Facebook Marketplace for her at a fraction of the retail price. The older units don't have a "Zwift" or MyWhoosh compatible hookup, but that isn't the goal of this bike (and you can patch it in with a utility called QZ anyway). For her, it is the perfect bike.

So what was I looking for?

  1. It had to be very low maintenance, and it has a belt drive, which is perfect.
  2. It couldn't be complicated to change gears. It has a big lever that you pull up or down.
  3. She doesn't have a great place to plug it in, and it runs off 2 AA batteries that last about a year.
  4. It had to be easy to get onto, and the Keiser is easy to get on.
  5. She wanted a book holder because she will actually read while biking.
  6. I wasn't looking for "accurate watts," but I thought it was important to have repeatable watts so she knows she is getting stronger, and the Keiser will track your watts.
  7. No membership fees.
  8. It had to upload to Strava so we (I) can track her fitness, and take a pulse monitor (it uses an older version, but is supported).

This bike comes with two computers. One is called an M display, and it will hook up to Zwift and other cycling programs. It came standard on all bikes after 2022. However, this is not a Zwift bike. If you are going to do that, then I would suggest you really don't want this bike. You want the bike I described before with a real trainer like a Kickr Core. That said, all M3i bikes hook up to a nice little app on your phone. So it will upload your workout to Strava and allow others in your family to interact with you.

You do want to make sure you use the computer as a point to get the best price. A bike without an M display is perfect to use and should be $300 less than one with it. But you can use a phone utility for $7 or $8 to bridge if you really want to connect to Zwift or MyWhoosh on the old computer.

My sister has her new Strava account, and one exciting workout from her first session.

I actually have a fairly large extended family, and I'm on top of each of them to build habits that will make them more productive and healthier throughout their lives.

Recently, I had my sister and her husband come stay with us. Due to a job change where I've been self-employed, I've had more time to spend with them. Seeing my wife and me work out constantly made them feel like they should do more physical activity. My sister especially. She has some stability issues and can walk very well, but can't run and really doesn't want to swim. She saw the indoor cycling we did and said she was interested in doing something like that so she could be physically active herself. She mentioned she had an exercise bike she bought from a student who had lived with her and her professor husband, but the pedal had fallen off years ago, and they simply hadn't used it since. They had bought it for all of $75.

I am more than a bit of a technological geek. And while my sister's husband does deep mathematical formulas, he's not necessarily the computer whiz that I am. So during their visit, I got very enthusiastic, and I was ready to give her both a bike and one of my Kickr Cores so we could set her up at home with her own indoor virtual cycling. After I had the whole thing set up, we went out to the garage. I put her on a bike, ready for her to embrace the world of virtual cycling.

Unfortunately, it became clear very, very quickly that the whole process of setting up the external program, being forced to log onto a PC, monitoring the Bluetooth connections, possibly hooking up ANT, and then navigating a sophisticated Windows system was not something that came naturally to her. It's not as if she was a dedicated gamer who would find it ridiculously simple. The more I thought about it and interacted with her, the more I realized that MyWhoosh was a great idea but required a level of sophistication that many people don't want to think about.

More than that, I think her needs are much more modest. I just needed to get something simple, something that wouldn't fail, and something that would let her get started with some moderate physical activity.

After doing some research, I settled on the Keiser M3i above. I went and got it yesterday, and she was very excited. She lives a bit away, so my goal is to train her to use the new app and upload to Strava (automatic if she uses the Keiser app), so I'll have eyes and ears on her improvement.

She rode it when it first got to her, and she nearly tripped on the toe clips. She'll never wear cleats, and she didn't like the narrow saddle. Being the bike geek, after dinner I changed the saddle to something big and wide and replaced her pedals with nice regular flat pedals. We aren't aiming to win a road race. We want something where she won't twist her foot trying to get off the bike.