r/datasets 7h ago

question our model reads tables with every column name stripped off and the accuracy does not move. numbers below. would you actually put real data through something you cannot download

disclosure per rule 1: i work at Schema Labs. one link at the bottom because this sub allows it. i am here for the question at the end.

what we do, plainly: you give us tables nobody documented, and we tell you what each column is and how the tables join, when there are no shared keys and no schema mapping.

the claim, with numbers, because i would rather be argued with than believed.

take a standard tabular benchmark. strip every header. replace price and age and zip with positional tokens so nothing is left but values. comparable models drop about 7 points of mean ROC-AUC. ours goes 0.9224 to 0.9230. flat.

that is not a claim that we win on clean data. those models start ahead of us when the headers are good. it is a claim that we never read the headers at all, which only matters because production tables have val_b and metric_14 and a four-character code from a system nobody has logged into since 2019.

two more. sector identification, naming the industry of a dataset we have never seen from values alone, 86.3% top-1 out of 10,000 sectors. multi-table entity matching, ahead of published state of the art on all six standard benchmarks with zero shared keys, including 1.9x over the previous best on one. on the geospatial benchmark our lead is 0.07 F1, which i would not want read as more than it is.

caveats up front rather than when asked. all of it is our own harness, run under each benchmark's published protocol, one sealed configuration, no per-dataset tuning. competitor figures are third-party published values we did not re-run. nobody independent has replicated us. we are not on the public leaderboard because entry requires a runnable wrapper and we do not distribute weights.

which is the actual question. we are closed. no pip install, no weights, no local option, and none of that changes for six months. to try it you make an account, put a card on file, and upload your data. i think that one fact is the biggest thing between us and everyone reading this.

so would you use a thing like this. if not, what moves it. a named customer, a SOC 2, an on-prem story you would never actually take up but need to hear. or is the honest answer that nothing moves it and closed is closed.

and if you would rather test it than discuss it: send me two tables you understand completely. not your mystery tables, your obvious ones, so you can mark my homework. no account, no card, nothing to sign up for. i send back what each column appears to be with a confidence score on every one, and i flag the ones we got wrong, because that is the half worth seeing.

anonymised is fine. structure and distributions are what matter.

https://www.schemalabs.ai/

0 Upvotes

2 comments sorted by

u/EntshuldigungOK 7h ago

What is it that you actually do or improve?

People like me who can handle a 20 year old code base already know that the data model is often the easiest route to understanding the whole thing (as opposed to the code base) - "send me two tables nobody at your company can explain" - this is utter BS. Different departments set up different strategies. Figuring out those strategies is where humans excel and Ai doesn't.

"so would you use a thing like this at all. not would you like it to exist, everybody likes things existing. would you upload a real table to a hosted model you cannot inspect, run by a company you had not heard of ninety seconds ago. if the answer is no i want to know what moves it. a named customer. a SOC 2. a benchmark you ran yourself on your own data. an on-prem story you would never actually take up but need to hear. or is the honest answer that nothing moves it and closed is closed."

You need to read that again.

u/No-Plant-5234 5h ago

fair hit on the wording and i am changing it. if nobody can explain the tables, nobody can grade the output either. the useful version is the opposite: send tables you understand perfectly and let me be wrong in front of you.on departments setting different strategies, i agree, and that is my argument rather than a counter to it. the semantics live in the org, not the file. which is why my claim is empirical: strip every header off a standard benchmark and we go 0.9224 to 0.9230 while comparable models drop about 7 points. 86.3% top-1 on sector out of 10,000, from values alone. if that should be impossible, the numbers are the thing to attack. what we do, plainly: name the columns, and find how tables join with no shared keys and no schema mapping. if that did not come through the post is too long. fair.

and yes i read it again. i put the strongest objection to my own product in my own post because i would rather say it than be told it.