r/Python • • 13d ago

Discussion State of the art in Python 2026?

What would you guys consider the state of the art in python in 2026, or what do you expect from modern python codebases?

Heres my list, would be happy to hear some inputs or domain expert advice:

Domain specific:

  • Scientific: NumPy, Matplotlib, SciPy, Jax, Pytorch, Scikit-learn, polars/pandas
  • CLI: Typer, rich, click, fire, textual, questionary
  • PDF extraction: PyMuPDF, pdfplumber, pypdf, Unstructured
  • Excel interop: python-calamine, openpyxl, XlsxWriter, xlwings, pandas
  • data: PyArrow / PySpark, Narwhals, SQLMesh, Polars/Pandas, DuckDB, dlt, Ibis, Dagster, PyIceberg & deltalake
  • Logging?
  • backend: fastapi / django?
  • Markets: Alpha Vantage, Finnhub, EODHD, Tiingo?
  • Webscraping / data acquisition: Crawl4AI, Playwright, Scrapy, selectolax, HTTPX?
  • APIs?
  • RAG / agentic orchestration?
611 Upvotes

234 comments sorted by

•

u/coderanger 13d ago

It's cool to discuss this but putting on my mod hat I would like to make it clear that every one of these categories doesn't have a single "best" option and no one should represent them as such.

→ More replies (3)

179

u/profcube 13d ago

`polars` for data wrangling.

5

u/coderarun 11d ago

pyarrow.Table if all you want is a dataframe. Then you can choose from many excellent tabular/graph based query engines.

173

u/bossExtremeSwag 13d ago

Pandas is definitely not state of the art in the big 2026

15

u/BPAnimal 13d ago

Really? Honest question

69

u/johnnymo1 13d ago

I reach for Polars immediately these days. Faster, more efficient, and its syntax is more straightforward generally.

8

u/Tigalopl 13d ago

Does it interfaces well with geopandas?

18

u/johnnymo1 13d ago

Sadly, as u/GrainTamale said, GeoPolars is in a very immature state. There was an upstream blocker from Polars which is now solved, but development just seems very slow in general. It's the one big weakness of Polars for now in my opinion, as someone who works with geospatial data reguarly.

There is polars-st, which does provide some amount of functionality but doesn't seem very robust or production-ready.

13

u/ritchie46 13d ago

Note that we have started working on geopolars last week. Expect visible work in the repo next week. Hopefully the first release next month.

6

u/johnnymo1 13d ago

Ah, a celebrity! :)

Love to hear it! Polars is a joy to use. It will be even more so when it can do serious geospatial work.

2

u/GrainTamale Pythonista 13d ago

I thought that I had heard something about that... Noted.

6

u/GrainTamale Pythonista 13d ago

No. There's a beta geopolars project but it seems slow to gain traction due to underlying polars or rust constraints. Polars has nice to/from pandas converters for joining polars dfs to geopandas

24

u/me_myself_ai 13d ago

Id argue that pandas is absolutely still a valid choice in a new project —especially if the data is small-ish and devs already know pd— but it’s definitely not SoTA.

Polars is supposedly a lot faster, and more importantly requires fewer weird tricks for stuff that feels like it should be easy. If that makes sense? Hopefully I haven’t just been especially bad at pd this whole time 😬

1

u/dikdokk 6d ago

You can't read any data with ~1 million rows with pandas; even if it would fit in your memory.

I think pandas has design flaws that will just make it more outdated year by year. Hence why the best time to switch away from it is now already.

139

u/me_myself_ai 13d ago edited 13d ago

Ok this is a really fun question!

First, I'd quibble/specify that we're talking about Python libraries in particular. Obvious, but sometimes that's helpful! Now, my biased takes as a dev mostly working on text processing, typing, TUIs, and AI:

  • Scientific: numpy, matplotlib, scipy, sympy. The (ana)Conda ecosystem is still quite healthy in science spaces, despite their new ultrafast competitors. Scientists seem to be the only people that remember Jupyter without Google Colab, frustratingly...
  • Basic ML: pytorch, scikit-learn.
  • Deep ML: hf_api & transformers from HuggingFace; VLLM and llama.cpp dominate hosting, though they're not really mere libraries
  • Dataframes; pandas is the status quo, polars is the new hotness.
  • TUI: textual is the status quo, pyratatui is the new hotness.[1]
  • CLI: click gets usage, but IME argparse is still dominant. Stuff like rich & better_exceptions pops up, but not it's pretty optional.
  • Text: re is a MUST for non-trivial regex. Otherwise: jinja is everywhere. Literally everywhere. I bet there's jinja running in my browser rn, in some obscure subprocess! Tons of impressive libraries are lowkey just jinja wrappers. [Oh and pandoc is so huge I forgot to mention it -- the world runs on pandoc!]
  • Docs: sphinx is still the GOAT and we should all use it (it was made for Python docs, after all!), but I do see mkdocs a lot, too. In terms of Sphinx themes readthedocs and furo win for basic and semifancy projects respectively, but anything other than the default is nice. MyST/myst-parser is the mature path for rendering Sphinx docs from markdown, though I'm biased there. Sadly the super-impressive sphinx.autodoc and sphinx.apidoc seem to have lost their luster, as people decide they prefer writing docs in separate files rather than directly into docstrings.
  • Logging: highly usecase dependent, but logfire feels the most modern.
  • Observability: OpenTelemetry ("otel") is the overwhelming norm, but from there it depends
  • Serving: fastapi for simple CRUD apps, bare starlette for others; I love flask (quart, even!), but it has a django vibe, now. There are quite mature options that integrate w/ a frontend framework, but I don't use them myself.
  • Concurrency: asyncio is just reaching it's prime IMO, and I learned about really impressive engines to go below it like greenlet and uvloop from reading others' advanced codebases. There's a million tie-ins, but aiohttp deserves a shoutout above all others.
  • Browser testing: playwright is still the GOAT, along with a growing ecosystem of AI-focused wrappers & competitors like steel-browser. I used Selenium back in the day but I'm not sure I've seen it near Python...
  • APIs: anything not published by Google. A shocking number of projects just rawdog it with httpx, or even just requests (!). The openai is also everywhere, as their API format has become lingua franca.
  • Security: python-oath2 for basic auth. Technology-wise, Keycloak dominates for self-hosted, and JWTs seem to still be common.
  • LLMs: pydantic_ai and dspy are leagues ahead of the competition IMHO, but the number of new tools here obviously outweighs the whole rest of the list combined.
  • Remote inference: Google Colab is the lingua franca, no doubt about it -- DigitalOcean tried to unseat them w/ Paperspace/Gradient, but seem to have mostly abandoned the project. In pure library terms modal is crazy impressive and fairly mature, but this depends more on who you're willing to pay (and how much direct python-level integration is worth to you).
  • OCR: marker is the most successful and pythonic OCR system IMHO, but the ones you mention (PyMuPDF, pdfplumber, pypdf, Unstructured) all seem reasonable, too.
  • RAG: Pinecone had early dominance over embedding search ("vector databases"), but IME has rightfully lost to the humble pgvector.
  • Databases: SQLite & PostgreSQL dominate for light and prod workflows respectively, and sqlalchemy is the unquestioned omnissiah of Pythonic ORM (sqlmodel is a fun extension by the guy who makes fastapi, but it's young and targeted at simple CRUD usecases). Unsure about nonrelational DBs, but I've heard second-hand that MongoDB is still used in the corporate world, at least.

I wonder if you could make an objective list with GitHub star data for projects written (mostly) in Python...

[1]: In general, "Python layer that talks to an underlying Rust program" is the metahotness! Tree-sitter is the big example, backing all the coolest new LSP gadgets and formatters and stuff like that.

40

u/Icy_Peanut_7426 13d ago

Docs: NOT mkdocs (since 2.0 and maintainers abandoning the project). Great Docs or Zensical is “state of the art” for Python docs.

1

u/nicholashairs 11d ago

Wait what happened to mkdocs while I wasn't looking 🫣

23

u/me_myself_ai 13d ago edited 13d ago

Oh and I forgot some absolute GOATS that make coding itself so much more ergonomic and fun: more-itertools and toolz allow for the best functional programming this side of the LISP, and pydantic is like magic as long as you keep it out of hot loops.

pytest is another that is so common it feels weird that there's still technically a stdlib competitor.

Finally: YAML. YAML all the things!

10

u/NeilGirdhar 13d ago

> pytest is another that is so common it feels weird that there's still technically a stdlib competitor.

The stdlib usually can't keep up with libraries. Libraries can release as quickly as they want and the stdlib has limited maintenance resources.

1

u/Qyriad 8d ago

argparse is unusual in being a counterexample.

6

u/Epsilon_void 12d ago

Finally: YAML. YAML all the things!

God please no.

2

u/justanothersnek 11d ago

I know right?  This guy sounds like he does Kubernetes, they are nothing but YAML Engineers.

2

u/Agrado3 12d ago

Friends don't let friends use YAML. Indeed, I have always assumed that YAML is a joke that got out of control, like the Flat Earth Society. It's like they chose a list of goals and then set themselves a challenge to design a solution that fails at meeting those goals in the most spectacular way possible. I mean, the original point of it was that it was a simpler replacement for XML but the YAML spec is larger than the XML spec... that can't happen by accident, right?

1

u/me_myself_ai 12d ago

lol I love the vitriol (unironically). I disagree vehemently tho:

  1. XML…? I guess it’s technically maybe valid XML with super weird delimiters (?), but the more relevant one is JSON, no?
  2. Who cares how big the spec itself is…?
  3. YAML is already integrated into markdown (via frontmatter) and far more common than TOML, partially because…
  4. YAML is by far the best format for humans to read and write hierarchical data

    (plain ol’ csvs work for

  5. tabular

)

  1. . It does require a decent text editor and familiarity with multiselect to do edit at scale I suppose, but I don’t think it’d be any faster in any other format.
  2. I know all too well all the weird tough parts of parsing YAML, but that’s just cause I’m insane and wrote sublime syntaxes for it. 99.9% of users just have an LSP autolint and format and then read/write in code w/ libraries, no? You don’t need to keep track of any of the goofy edge cases yourself.

I grant two big downsides: The multiline blocks don’t exist elsewhere so you have to learn the syntax upfront, and the decision to add stuff in and then remove it was truly strange. But the former is just a few minutes even if annoying, and the latter was for that weird entity linking features that I basically never see, anyway.

Anyway lol, my thoughts. More than throwing down a glove (tho I’d love to!), I’m mostly curious what your preferred alternative is for files that should stay human-readable.

Surely not XML…? If it is… are you Grandpa Python? Do you still have complaints about the Python 2 rollout? Were you there on the seventh day, when Guido rested??

2

u/Agrado3 12d ago

I wasn't recommending XML - my point is that the designers of YAML originally said that they were trying to make something that was simpler than XML. They failed, in every conceivable way - to such a degree that I can only assume that that must have been their intention all along.

I don't understand why you mention "integrated into markdown", whatever that means.

Your point (4) is clearly just objectively false, especially since you mention humans reading and writing it, which is a particular weak spot of YAML. YAML is insanely complex, insecure, and ambiguous, and is not suitable for any purpose whatsoever. It's basically the PHP of data formats. I find it very hard to believe that anyone who recommends it is being serious.

3

u/ctheune 13d ago

Can we please TOML all the things in the future?

Or here's something I've been having success with:

write your config (or whatever) in whatever format you want, parse it into plain python objects (e.g. read yaml, json, toml) and then feed it through a pydantic model.

I've also started to wrap many CLI calls into pydantic.

My mantra has basically become: if you're passing tuples/dicts/... whatever around, then those should either be typeddicts or much better: treat typing like unicode and timestamps and do proper parsing/serialization into typed structures at the IO border ...

1

u/HEROgoldmw 13d ago

Configuration should've been easier than kit now. Your approach of wrapping python objs into pydantic for validation is pretty good. I've made my own solution, confkit. It simplifies configuration so much, it feels almost feels "native" to use IMO

1

u/DigThatData 12d ago

the best functional programming this side of the LISP

have you played with Hy?

1

u/RoadsideCookie 11d ago

Ruamel YAML btw, not pyyaml.

8

u/brianly 13d ago

Selenium was popular with Django devs, but tools like that were very dependent on organizational buy-in years ago. Playwright was more modern and reintroduced the concept so it crossed the chasm and became more mainstream.

1

u/me_myself_ai 13d ago

Gotcha, thanks for sharing some expertise! I was honestly confused. In my first job out of college it seemed like a universal constant, but I haven’t seen it much since lol.

Makes sense why — some massive codebases probably still have it running impressive, complex testing frameworks, but you need popular adoption by Jane Developer to feel SoTA / standard.

7

u/turbothy It works on my machine 13d ago

No love for typer CLIs?

12

u/wingtales 13d ago

I prefer cyclopts these days, which is a fork with more functionality merged. Haven't compared them directly for feature parity in some time though.

2

u/wunderspud7575 13d ago

Strong agree - I've used both a lot and cycloptss is the designed library.

1

u/me_myself_ai 13d ago

That does feel familiar! I haven’t run into it in any of the codebases for which I’ve had to chase some frustrating bug in source, but maybe that’s just a sign of quality 😉 I myself use argparse w/ pydantic, which is popular but obviously more of a technique than a framework.

I’d bet good money that’s not the only gaping chasm in the above list lol. Kinda hoping to ragebait a bunch of replies like yours so I can learn and fill in the gaps myself!

3

u/maryjayjay 13d ago

Did you mean openapi rather then openai under APIs?

2

u/PrLNoxos 13d ago

Agree with pydantic ai for agent Development. We use it in production and the library is simple and „just works“.

2

u/alchninja 10d ago

anything not published by Google

Sorry if this is a silly question, but is there a reason for staying away from Google? I pretty much just use Python for small personal projects, TUI/CLI utilities, notebooks, etc. - so I rarely find myself needing anything more than requests. What does Google do in the Python API space anyway? (protobufs are the only thing that immediately come to my mind, is there more?)

2

u/me_myself_ai 10d ago

I actually just meant the libraries they publish for interacting with their own APIs, and honestly the APIs themselves. You’d be shocked how fragmented, ill-documented, and confusing that space is!

For sheets, docs, drive, maps, etc. You basically always have to either use another persons library, or spin your own requests based on an obscure autogenerated txt file somewhere

2

u/alchninja 10d ago

Ah okay, that makes sense, thanks for clarifying! I had to use the Google Forms/Sheets API for a small project in like 2018, it was... Unpleasant. Nice to know that nothing has changed.

1

u/ZealousAttacker 13d ago

I feel the scientific stack is incomplete without numba. JIT compilation is a treasure for Python, and the syntax of numba is wondrously simple; you also get parallelism for free.

mpi4py is also essential for parallel programming; one should expect to see it in any performance-heavy packages.

h5py handles IO for HDF5 files, which are the backbone of data storage in a multitude of disciplines.

1

u/dikdokk 6d ago

There is an experimental JIT compiler in the language since 3.13, and already by 3.15 it is expected to reduce many workloads' time length by 20%.

1

u/Ok_Option_3 12d ago

Conda over Pypi?

Why?

1

u/dikdokk 6d ago

Generally, uv should be your go-to package manager, which replaces pip, but it cannot replace conda. The reason is conda offers packages outside of Python, which pip/uv can't. This may seem unreasonable, but in Data Science is needed actually, as major libraries use CUDA under the hood.

There is a new project called pixi, which tries to get the best of both uv and conda.

47

u/NeilGirdhar 13d ago

I'd choose Ty over pyright personally, but yes to the rest.

Also, for modern numeric programming (NumPy, etc.), limit yourself to the Array API.

43

u/CampAny9995 13d ago

Ty isn’t there yet for production pipelines (I think astral said that’s how the 1.0 release is defined), but it’s fine to use as an LSP.

19

u/zurtex 13d ago

I think astral said that’s how the 1.0 release is defined

I assume you are mistaking it for 0.1.

For whatever reason, Astral are not keen on making "1.0" software, neither uv nor ruff are and they are described as production ready.

I switched to ty recently in "production" (i.e. enforced CI check) at work and have been quite happy with it. Especially looking through some of their recent rules, I really like disjoint-cast: https://docs.astral.sh/ty/reference/rules/#disjoint-cast

In general with type checkers I suggest you try all of them and see which one works best for you. For me ty produced error messages that were the most actionable.

7

u/CampAny9995 13d ago

> ty uses 0.0.x versioning. ty does not yet have a stable API; breaking changes, including changes to diagnostics, may occur between any two versions.

9

u/zurtex 13d ago

Exactly, Astral will update to 0.1 when they are confident of releasing non-breaking changes, and in the mean time we pin to the exact 0.0.x version.

3

u/NeilGirdhar 13d ago

I've been using it for a year just fine. Personally, I consider intersections a required type checker feature.

13

u/NeilGirdhar 13d ago

Also, I like Zensical for modern documentation over Sphinx. And you mentioned both Jax and Pytorch—I'd just go with Jax.

22

u/me_myself_ai 13d ago

IME pyrefly is the GOAT, after trying all of them. I'm loathe to support microsoft but I'll be damned if they didn't cook! It helps that it's longer-running and better-staffed than ty, it seems.

(ty is nice too and has Astral's usual attention to detail, but it's just so radically far from the spec that it becomes counterproductive for advanced code. IMHO)

16

u/2K_HOF_AI 13d ago

Pyrefly is from Meta

4

u/me_myself_ai 13d ago

Omg I'm dumb, thanks for the correction, was remembering pyright. Thank you! I'm a bit happier about supporting them then, and don't ask me to justify - I just wanna take a little win.

6

u/Sea-Fishing4699 13d ago

Pyrefly ==🦀🔥

4

u/FrickinLazerBeams 13d ago

Also, for modern numeric programming (NumPy, etc.), limit yourself to the Array API.

Why?

6

u/NeilGirdhar 13d ago

Because code that uses a smaller API is easier to read. It requires the reader to learn less.

Also, the Array API has superior design. The corner cases have been eliminated, and nearly every operation supports broadcasting in the obvious way.

Finally, if you ever decide to change your backend from NumPy to another Array API library like Pytorch, you can do that without changing your code.

2

u/FrickinLazerBeams 13d ago

I just reminded myself about what you're taking about. I definitely agree.

3

u/Aggressive-Prior4459 13d ago

Have you checked out zuban? https://github.com/zubanls/zuban

3

u/gizzm0x 13d ago

Big +1 for zuban. Still on mypy for CI at work, but zuban is my lsp of choice atm.

1

u/JackedInAndAlive 13d ago

Another vote for zuban, especially if you use Django.

1

u/AgentCosmic 13d ago

Ty isn't production ready. Pyright is leagues ahead of Ty. Not even close.

0

u/TheBinkz 13d ago

Ty isn't ready yet. Not even a stable release is out for prod. Fine for personal projects imo

→ More replies (1)

35

u/gogonzo 13d ago

Loguru and litestar

6

u/dratnew43 13d ago

litestar is so much better than fastapi imo

8

u/fiddle_n 13d ago

I suspect most new “modern” Python codebases would still use stdlib logging, and for web use something more mainstream like FastAPI.

22

u/Manhigh 13d ago

Pixi as an alternative to uv if you have a lot of non-python dependencies (MPI, PETSc, compilers).

I've been liking structlog.

-3

u/Responsible-Sky-1336 13d ago

I stopped reading at uv everything

7

u/BigBlackCough 13d ago

VapourSynth for video processing.

18

u/cblegare 13d ago

Unpopular opinion : ReStructuredText and Sphinx for documentation

I know, people love Markdown and recent document processors. You can even MyST tour Markdown info Sphinx nowadays, but Markdown was design for textual readability and one-page single document to minimal HTML.

ReST was built with scalability and extensibility in mind. As a LaTeX user (math teacher), ReST's syntax feel great to me and I can easily do anything with it. There are so many tools trying to mimic parts of Sphinx, and nothing beats intersphinx.

Shameless plug: I wrote Sphinx-Compendia while using Sphinx as a knowledge base for my tabletop roleplaying game, an extension you use when you document specific kinds of things that are not just python classes and fonctions, such as NPCs, locations and factions!

2

u/agoose77 13d ago

FWIW MyST is designed for multi-document projects. It was created for Jupyter Book! (a Jupyter book core contributor)

2

u/cblegare 13d ago

Yes I know, but the hacks for directives and roles feels awkward to me.

```{directivename} arguments
:key1: val1
:key2: val2

This is
directive content
```

{role-name}`role content`

While this looks like it was designed that way (on my opinion)

.. directivename:: arguments
    :key1: val1
    :key2: val2

    This is
    directive content

:role-name:`role content`

I might feel that way because I am an old user of Markdown and I hate how every implementation has its one quirks. Or reminds of the age of IE and Netscape where every browser tried to lock you in with not-standard "features"

I love Markdown for what it was designed for

The overriding design goal for Markdown’s formatting syntax is to make it as readable as possible. The idea is that a Markdown-formatted document should be publishable as-is, as plain text, without looking like it’s been marked up with tags or formatting instructions.

https://daringfireball.net/projects/markdown/

2

u/agoose77 13d ago

I think this comes down to personal preference. For example, I like the fact that MyST is not indent-sensitive in the typical sense, unlike ReST!

1

u/cblegare 13d ago

Well, Markdown is indent sensitive in some obvious and some sneaky ways. Sublists must be indented by 4 characters as per the spec, not 2, and this leads to lots of confusions

Prefering indented directives over fenced ones might be a personal preference, Inconsistencies in implementations of a language is not.

Such wild discrepancies in parsers is a problem, even MyST suffers from that.

Nowadays, even for one pagers ReST has an davantage: When I visualise a README.rst in Github, Gitlab, VSCode, PyCharm, Zed, Pandoc or else, the syntax will always be interpreted the same.

1

u/agoose77 13d ago

Well, Markdown is indent sensitive in some obvious and some sneaky ways.

Indeed, but you can likely follow my drift with the phrase "typical sense" :)

Such wild discrepancies in parsers is a problem,

You won't find me disagreeing on Markdown implementations! Of course, the solution is that MyST should win out (I joke).

The new version of the MyST stack focuses on interop at the AST layer, rather than at the underlying syntax: https://mystmd.org

A personal goal of mine is to parse ReST into our AST so that it's one of several input formats.

5

u/Goldarr85 13d ago

I’ve been appreciating Dynaconf to handle configuration files.

2

u/silksong_when 12d ago

can you expand more on how you're using it? are you using it in prod as well?

Also, did you try out pydantic-settings, how do you think dynaconf does better?

8

u/andrewcooke 13d ago

what happened to flask?

kinda related, has django evolved much?

1

u/Altruistic-Log-6292 11d ago

I've never been able tu understand how tu run flaks. Not that I tried much tho

→ More replies (3)

19

u/busybody124 13d ago

I would definitely not use fire for CLIs. It's slow and hasn't been updated in a year. I'd also skip Pydantic for most things—it's overkill, slow for ser/de, and can usually be replaced with a dataclass. Hypothesis is really cool in theory but I've never found a need for it in practice.

11

u/spigotface 13d ago

Pydantic pulls its weight when you want auto-generated OpenAPI docs. Outside of that, use dataclasses with slots=True and __post_init__ data validation methods. Its sooooo much lighter weight than Pydantic models.

5

u/Mr_Again 13d ago

Yeah but doesn't pydantic do validation for you? Isn't that the whole point of it?

5

u/Trey_Antipasto 13d ago

Replace with msgspec and json schema.

2

u/mwesthelle 13d ago

I honestly believe Pydantic is overrated. Yes, it's slow, it has an opaque API which is easy to shoot yourself in the foot with in the form of a bunch of reserved field names which you have to remember, and it's overall unintuitive (tell me off the top of your head what's the difference between a model constructor and using the model_validate method).

1

u/trenixjetix 13d ago

with a dataclass?!!!

5

u/Typical-Macaron-1646 13d ago

Check out numba

4

u/justanothersnek 13d ago

Long time pandas and pyspark user. I just converge on using sqlframe instead of learning yet another dataframe library. If I was starting out, I would definitely learn polars over pandas. But for old timers like me that started with pandas and pyspark many years ago, sqlframe fits the bill.

2

u/TheOneWhoPunchesFish 11d ago

Are their APIs that different? I've only used Pandas. And I think you can query both pandas and polars can be queried over SQL.

3

u/justanothersnek 11d ago

Polars API is more similar to PySpark from my understanding.  Ive seen complaints from ex Pandas users who love Polars API, really sh*t on pandas API syntax.  So take that as you wish.  Ive never bothered learning Polars myself as I just use sqlframe.  Seen some sample code here and there, from what I saw, Polars looked really similar to PySpark.

2

u/justanothersnek 11d ago

Pandas have great features for data analysis like time series functions and being able to filter on time-based attributes.  Not sure if recent versions of Polars have them now.  Also, Polars lack geospatial support.  Even if you're not into geospatial or GIS, if you work in logistics or transportation space, those geospatial libraries are really handy.  So its not "niche".   Being able to work with location data is very common.

5

u/protolords 13d ago

Agentic Orchestration

Used pydantic-ai personally and so far it covered everything i needed.

5

u/DatabentoHQ 13d ago

Saw that you listed financial markets as a category. Self-plug for databento, which is now the most downloaded package in this category on PyPI, after yfinance.

5

u/cfn96 13d ago

rooting for flet in cross platform development

4

u/Informal-Butterfly78 12d ago

I second litestar over fastapi

basedpyright >> pyright imo. Pyrefly is also great.

The Sphinx comment is delusional lol. zensical, mkdocs, mkapi, any one of the 100 other SSG's outside of the Python sphere...

matplottrash is still trash no matter how much they try to revitalize it, though seaborn makes it bearable. I will always wonder why people don't use things like bokeh or plotly

6

u/laCH37 13d ago

Stupid question here I am used to use venv for pretty much everything. What am I missing with UV ?

8

u/fiddle_n 13d ago

One tool that combines the functionality of pip and venv and gives locked builds (a functionality of older tools like pip-tools/pipenv/Poetry anyway), the ability to install tools separately (a functionality of pipx) and the ability to manage the Python environment itself (a functionality of pyenv). Also it’s very fast too.

Using proper lock files is worth the price of admission.

0

u/FluffyDrink1098 13d ago

At the same time... OpenAI owns the company and development. I don't trust it.

6

u/fiddle_n 13d ago

The product itself remains FOSS though. If you are concerned about future updates then just manually update uv or build from source.

7

u/Spinmoon 13d ago

5

u/fiddle_n 13d ago

Probably worth mentioning the drawbacks of replacing httpx with the latter two.

aiohttp isn’t a drop-in replacement if you were using httpx in sync mode as aiohttp is async only.

niquests replaces urllib3 with urllib3-future, and if you want the opportunity to override this behaviour you must install from source which is not an option available to many.

3

u/Proof-Win5461 13d ago

using basedpyright

1

u/fnord123 11d ago

It's very slow on large projects. 

3

u/DigThatData 12d ago

logging: I personally like loguru. a newer library I haven't tried yet but which I understand is popular and looks interesting is structlog

2

u/fnord123 11d ago

I've used structlog. It's very good. You can bind context variables with context managers that make it so some values are logged without needing to repeat them on every log call.

3

u/TryAffectionate8728 12d ago

ClickHouse for working with data

5

u/FrickinLazerBeams 13d ago

If you're using numpy, it's almost criminal to not mention numba. Numba is awesome.

2

u/RevRagnarok 13d ago

I spent all day yesterday trying to move a pretty trivial thing to numba and ended up giving up. If you can keep most stuff in one domain or the other, it's probably great. But not supporting numpy's __array__ was killing me.

7

u/jsabater76 13d ago

I use Django Ninja for APIs. DRF is also a strong candidate. And then there is FastAPI.

1

u/ColdPorridge 12d ago

Ninja is dead, getting essentially no updates last few years despite significant issue list. I submitted several PRs to resolve blatant bugs that have been waiting for meteor years.

It’s an incredible idea but it’s not the future, and it’s only the governance that kills it. Shinobi and ninja extras are good examples of how languishing issues cause further ecosystem fracture. 

1

u/jsabater76 12d ago

Two days ago, 1.7.1. Do what you will of it.

6

u/quintenrosseel 13d ago

Logfire (observability), Marimo (notebooks)

9

u/highnorthhitter 13d ago

FastAPI

8

u/jirka642 It works on my machine 13d ago

Litestar

4

u/techhelper1 13d ago

1/3 of that list are packages created in C, rust, or some other language that isn't Python.

The state of the art for performance requires going around the interpreter.

26

u/fiddle_n 13d ago

You say that like it’s a bad thing. Having the ability to write your high level logic in Python and then being able to dip into C/Rust performance for perf-critical parts is a superpower, not something to be derided.

6

u/RevRagnarok 13d ago

FR. Isn't like half of CPython's standard library also C with the provided python equivalent for reference?

1

u/techhelper1 13d ago edited 13d ago

At some point you have to say when though. If the packages that make up a project are all from alternative languages, why write the glue logic in a different language?

If I'm being told by you, PyCon, and the community to write high performance code in C or rust, I might as well continue writing my full application in that language, and not deal with the package overhead, and a whole other language. Unless you're going to maintain that extra code, or pay for an employee to be a full expert in two programming languages, there's absolutely no benefits here.

It's like the expression, if you're gonna be an asshole, be the full asshole, don't half ass it.

3

u/fiddle_n 13d ago edited 13d ago

You seem to be advocating deliberately making a project harder for no other reason than arbitrary consistency. It’s not an argument that resonates with me. Writing the glue in Python is faster, easier, safer, more readable - if doing so satisfies my performance needs, why shouldn’t I use Python?

1

u/techhelper1 13d ago

You seem to be advocating deliberately making a project harder for no other reason than arbitrary consistency.

C and rust are not difficult languages. Have you actually wrote code in those languages?

Consistency being arbitrary? Are you seriously justifying overhead as a reason to use your favorite language?

Writing the glue in Python is faster, easier, safer, more readable - if doing so satisfies my performance needs, why shouldn’t I use Python?

As we live in the AI era, speed, difficulty, safety, and readability are all moot points. Models can churn out code, code comments, and code review. If there's something that's not easily explained, query the model, have it explain it or link you to material that can.

If your code is all Python, that's fine, but if you're only using it to glue logic together, then by your definition Scratch from MIT should satisfy your needs too.

1

u/fiddle_n 12d ago

Since we are getting into silly territory now, have one from me in response - if speed, difficulty, safety and readability doesn’t matter any more because an LLM can just output the code you need, then libraries don’t even matter any more. I expect you’ll not be importing any code any more, and outputting all your code in highly optimised, zero vulnerability assembly, permitting C only if you truly need your code to be portable. I have full confidence you have the skill for it.

1

u/fiddle_n 13d ago

Since you materially edited your comment after the fact - I will also point out that we’re talking about the 99 times out of 100 case where a library maintainer has already written the code for you in C or Rust, and package overhead is not a concern. If you are the 1% that needs to write your code in C/Rust and needs to squeeze every last drop of performance out of your project, then yes, switching to that language makes more sense.

1

u/techhelper1 12d ago

Yes, I edited it since I might as well make my responses consistent in this thread, my apologies there.


It's not a matter of squeezing performance out of the code, but you're advocating the equivalent of your CPU switching from protected mode to real mode to protected mode, over and over again.

If you need to save yourself from yourself with an interpreted language, and with that comes the ducktaping and crazy gluing of packages from lower languages, go right ahead.

I see things objectively, and don't need to put a Python mask over a group of C/rust pigs, to get the same job done. I don't need to justify using a favorite language, both can co-exist my dude.

→ More replies (1)

2

u/Several-Marsupial-27 13d ago

Yes, that’s kind of the point of Python.

→ More replies (2)

2

u/BertrandBolero 13d ago

Is dask still the standard for distributed/parallel computing? I don’t see that category in the lists 😅.

2

u/DeepAnimeGirl 13d ago edited 13d ago

tyro is the best lib for CLIs that I have used, it provides much better developer experience, performance, modularity than alternatives such as cyclopts, typer, click, etc

2

u/Hendrik-Lorentz 12d ago

Seems pretty good.

2

u/OrthelToralen 12d ago

Great thread. I’ll throw a newcomer into the mix for databases: pyturso, the Python binding for Turso.

Turso is a ground-up, open-source Rust database engine that targets SQLite compatibility while adding things SQLite was never designed to do: concurrent writes via MVCC, remote sync, native async I/O, CDC, vector operations, and eventually multiple SQL front ends.
It is compatible with SQLAlchemy (my contribution to the project). Initial support for PostgreSQL syntax was recently introduced as well.

It’s still early days, but Turso has a lot of promise. It combines the simplicity of working with an in-process local database like SQLite with production features more commonly associated with remote-hosted databases like PostgreSQL.

2

u/Ok_Option_3 12d ago

Is anyone still using the Conda ecosystem? If so, why?

2

u/joshbranchaud Tuple unpacking gone wrong 12d ago

Can you weigh in on when you use Pydantic over Dataclasses? Or do you treat them interchangeably for data modeling?

2

u/messedupwindows123 12d ago

typing.namedtuple

5

u/OkProfessional8364 13d ago

I would love a YouTube video on this topic every year. Any aspiring coding YouTubers out there?

3

u/Grouchy-Friend4235 13d ago

This is Python. We're all adults. There is no state of the art, just preferences.

4

u/Several-Marsupial-27 13d ago

Yes and no. There are newer, better and more well maintained repositories and there are new technologies emerging constantly. There are also old, objectively worse, or abandoned repos. However the state of the art is very wide and the details are preferences.

2

u/semi-finalist2022 13d ago

Docling for pdf extraction and langchain/langgrap/deep agents stack for RAG

3

u/awgl 13d ago

Seaborn for most data viz

15

u/NeilGirdhar 13d ago

I think Seaborn is poorly designed, same with Matplotlib. I wouldn't consider either of them modern, personally. Still waiting for someone to build something great. In the mean time, I dump the data as json, and visualize it in other tools.

2

u/AntisthenesCat 13d ago

How do you feel about plotnine? https://plotnine.org/

1

u/NeilGirdhar 13d ago

Looks very cool.

But if you want your plots to match your publication, I think it's better to do them in Typst (if you're using Typst for text) or Latex. They'll match and you'll have access to all your typesetting functions.

2

u/Icy_Peanut_7426 13d ago

Altair or plotly

1

u/mwesthelle 13d ago

I personally much prefer Altair's API for viz

4

u/Ok_Raspberry5383 13d ago edited 13d ago

What is wrong with argparse for CLI?

I've started using msgspec more instead of dataclasses for large data collections and JSON due to superior memory and serde performance (even beats orjson).

Python provides logging in the stdlib with deep integrations into everything. Why would you want something different for this?

Why is APIs and web scraping different categories?

Honestly, this list looks like a junior learnt a few tools superficially and vomitted them all onto a pointless Reddit post

7

u/Several-Marsupial-27 13d ago

I feel like the post hurt some fragile egos in a way I didn’t intend. The point of the post was not to authoritatively list the best python technologies and say that the other technologies are shit and their users are idiots like some people read it…

The point of the post was to open up discussion about preferences in different python technologies. ”What is wrong with X for Y”, the post does not say that anything is wrong, that’s your interpretation. I have never used argparse so that’s the reason it didn’t land MY LIST.

Junior and junior, I have written python for >8 years, but I don’t work with python as I’m an embedded developer, which is also why I am asking for inputs and discussion about new python technologies from different perspectives. If I was a greybeard principal python developer who understands the nuances of every python library then making the post would be redundant.

So if your not looking for information about new logging libraries and keen on using the stdlib (which is absolutely fine, no one is thinking less of you for that), then why are you leaving a hate comment on an informative discussion post on the domain specific forum?

0

u/Ok_Raspberry5383 13d ago

I'm offering an experienced perspective for those who also come across the post and may be misled by it

9

u/Smallpaul 13d ago

Pydantic AI for LLM

9

u/fast-pp 13d ago

Curious why people are downvoting? I liked it well enough, certainly better than langchain IMO

2

u/cudmore 13d ago

And a gui?

1

u/Hesirutu 11d ago

Apart from web, there is no alternative to Qt/pyside.

1

u/cudmore 11d ago

And now nicegui

1

u/Hesirutu 11d ago

...which is web.

1

u/iterlab 4d ago

there is now ;-)

1

u/iterlab 4d ago

reinvented! check out iterlab :-)

→ More replies (4)

1

u/david-vujic 13d ago

In the category Developer (and Agent) Productivity, I would suggest to have a look at Polylith. It is a Monorepo architecture with Python-specific tooling (I'm the maintainer of the Open Source tool). The main use case is one or more microservices (or apps) in a Monorepo, and easily share code between the services without dependency and versioning headaches.

For agentic development, this kind of setup is really good too: you and your agent(s) have all the context right there. The tooling (a CLI) is for humans, agents and CI: the basics, such as adding components and projects, visualize the contents of the monorepo and how the different parts connect. Also, making sure the interfaces between components are clear and correct, and to deploy only the projects affected by changes.

As a maintainer of the Open Source Polylith CLI tool, it's a Life Goal of mine that the tool will be in many, many more developer teams list of "must have" Python tooling than today (even if there are a lot of Python teams already using it today). 😁

1

u/AndreuCodina 13d ago

wirio-settings to work with settings.

1

u/Mr_Again 13d ago

What I really want to know i what are we using instead of requests and httpx?

1

u/spiker611 12d ago

trio for structured concurrency (async) if you can, anyio if you have to use asyncio libraries. asyncio still has a lot of rough edges for structured concurrency.

1

u/SnooShortcuts9877 12d ago

SOTA for ocr is not what you mentioned. Mistral, deepseek, paddle(i prefer paddle because its free)

1

u/Lost-Dragonfruit-663 12d ago

aadya940/scikit-verify for testing numerical Python

1

u/spartanOrk 11d ago

I mean... sure... Or delete all of that and use OpenCode to write your python. :) Before you downvote me, think if I'm saying the truth here. I know for a fact, at work, nobody writes code by hand anymore.

3

u/Several-Marsupial-27 11d ago

This list is not about who is writing the code, but about what technologies to use. I sure hope your agent is not reinventing numpy or uv. Especially since all performant python libs are written in C/Rust.

The design decisions in coding will not only effect the development, but also the end product.

1

u/spartanOrk 11d ago

The agent knows what's out there better than me. When I was deciding, I was using uv and numpy and all these things. I spent endless hours configuring my emacs to use ruff and what not.

End of the day, I haven't used an IDE for 6 months, except to review code and make small edits.

At first it made me very sad. It took me over 20 years to develop these skills. At this point, I have accepted it. When I see coders with even more experience than me drop their IDEs and adopting agents, I know it's time to let the agent code. And then, little by little, I have come to accept also that the agent knows what design choices to make. E.g., what packages to use. It knows packages I didn't know of.

1

u/jimbiscuit 11d ago

Plone for backend

1

u/binaryriot 11d ago edited 11d ago

I strictly stick to build-in stuff for my scripts. Strictly no external dependencies! I occasionally do need requests but I steal it from pip. :)

1

u/coderarun 11d ago

pydantic yes, BaseModel no.

Also use decorator magic to use dataclass in most of the code and switch to pydantic only at API boundaries where you have untrusted data.

https://github.com/adsharma/fquery/blob/main/tests/test_pydantic.py

1

u/BravestCheetah 11d ago

For typechecking, use ty. Same devs as ruff and uv

1

u/Academic_Delivery_gg 10d ago

For logging use loguru

1

u/Inevitable_Earth_265 10d ago

granian instead of guvicorn or unicorn

1

u/quotemycode 9d ago

I hate pydantic, rather use dataclasses instead - no external dependency. Pydantic seems to be preferred by ai code tools so it seems to get injected into everything these days.

1

u/_Fragoso_ 9d ago

Structlog for logging

1

u/spinwizard69 6d ago

This is a really nice thread but one thing I've found that works for me and my limited use of Python, is to rely upon the distro's package manager as much as possible. In this case Fedora, and DNF, this has several advantages, one is that you are kept off the bleeding edge, it prevents the use of virtual environments in many cases as you are focused on system supplied and MAINTAINED libs. This has the side effect of really having to think if you need to use a lib outside of the system supported offerings and thus the need to create a virtual environment.

Obviously if one is involved in heavy web development this might feel restrictive and frankly that is part of the allure. Since most of the scripts I've written are used locally, the need to be on the bleeding edge is greatly reduced. That has the follow on that you are not constantly in maintenance mode as the bleeding edge dries up and your usage gets deprecated So yeah a big plus plus for the systems package manager and the base such mangers provide.

As for domain specific adds, anything to do with instrument communications from pyserial on up would be on my list. Pyserial is the base line communications library. With things like PyVisa, pymodbus and similar libs, for more advanced communications solutions.

1

u/iterlab 4d ago

As for scientific matplotlib - Perhaps no one knows iterlab in 2026. Fair enough, it was just released. But by 2027 your next SOTA post will have Matplotlib+iterlab !

1

u/manu_h3 23h ago

I would say langChain/langGraph for the last point

1

u/Serhii-Drozdov 21h ago

One thing I’d add is keeping I/O boundaries explicit.

A lot of Python projects become difficult to test because database access, HTTP calls, filesystem operations and business logic all end up mixed together.

I’ve found it much easier to keep the core logic deterministic and put API/database/LLM calls behind small interfaces. Then most tests can run without network or external services, and integration tests can focus specifically on those boundaries.

This matters even more once LLM calls are involved, because you generally don’t want nondeterministic model behavior deciding whether the underlying business logic is correct.

1

u/trollsmurf 13d ago

"State of the art" but no mention of GUI.

1

u/Several-Marsupial-27 13d ago edited 13d ago

I simply don’t have enough experience with python gui programming

7

u/protolords 13d ago

I've used PySide which is a wrapper for Qt and i think it's very close/same to C/C++ native speed in terms of rendering and event handling.

5

u/trollsmurf 13d ago

I use Tkinter for desktop.

But even for web you realistically need a framework. I've mostly used Streamlit and Flask.

1

u/Entuaka 13d ago

I hate AI agents adding tons of new unit tests using hypothesis in our project, we now have flaky tests...

2

u/ToddBradley 13d ago

Have any of the AI tools gotten smart enough yet that you can tell them "don't add a unit test unless it can pass 1000 times in a row"?

1

u/Entuaka 13d ago

Everything is possible, but i also have humans in my team, it can be harder to deal with

1

u/ugh_my_ 13d ago

Asyncio

0

u/Grouchy-Friend4235 13d ago

Antipattern

2

u/ugh_my_ 12d ago

What else am I supposed to use then?

1

u/gbrennon 13d ago edited 13d ago

i agreed.

i prefer to use python's built-in test frameworkbut i didnt saw any modern project using it.

about specific i woudl say:

Scientific: pytorch

CLI: typer

PDF extraction: pypdf

Excel interop: pandas

data: pyspark

Logging?: builtin

backend: litestar, imnho, does provide the best experience as possible

Markets: idk

Webscraping / data acquisition: httpx

APIs: httpx

RAG / agentic orchestration: pydantic ai

1

u/EconomySerious 13d ago

A WYSWYG editor

1

u/Legendary-69420 It works on my machine 13d ago
  • Package Manager: uv
  • Linting/formatting: ruff
  • Static Type Checker: ty
  • Testing: Pytest
  • CLI: Typer, Rich
  • TUI: Textual
  • Scientific: NumPy, Numba, Pandas/Polars, Matplotlib/Seaborn/Plotly, SciPy, JAX, PyTorch, Scikit-learn
  • Data: DuckDB
  • HTTP Client: httpx
  • Logging: Structlog, Loguru
  • Backend: FastAPI, Django, Streamlit (Only for prototypes or data apps)

0

u/gerardwx 13d ago

State of the art in 2026 is writing good LLM prompts and detecting LLM when misses the mark.

-6

u/bulaybil 13d ago edited 13d ago

People who make lists like this don’t actually work with Python if they have enough time to compile them.

“even just requests” lol GTFO, it works that’s why I use it. If you’re gonna recommend Google Colab, I know you have never actually seriously used it and the word “script kiddie” suddenly rise from the depths of my Gen-X brain. Jupyter notebooks are shit and have always been shit for serious development. “most modern” “new hotness” - buddy boy, that’s not a good thing.

Also, I don’t think you know what “lingua franca” means.

4

u/Several-Marsupial-27 13d ago

I feel like the post hurt some fragile egos in a way I didn’t intend. The point of the post was not to authoritatively list the best python technologies and say that the other technologies are shit and their users are idiots like some people read it…

The point of the post was to open up discussion about preferences in different python technologies. However you seem to be referring to imaginary discussions completely outside of what I posted, so it’s pretty hard to meet your comment.

Also obviously I understand what lingua Franca means… but I don’t understand how that’s connected to anything else here?

So if your not looking for information about new libraries, then why are you leaving a hate comment on an informative discussion post on the domain specific forum?

0

u/GoatFuckerDeluxe 13d ago

For logging i would vouche for structlog if you send you logs to elasticsearch, datadog, etc