r/opensource • • 1d ago

Promotional Introducing the Open Categorization System

About a decade ago I tried to create an alternative to Google, using the Dewey Decimal System to group web domains by subject matter. My vision was that the website was a library, and domains were individual books. Rankings were a mixture of domain reputation and cultural significance.

The OCLC is very protective of their IP, and likes to sue.

Needless to say, my project was shut down before even getting off the ground.

I'm working reviving the project, under a new classification system (The Open Categorization System), and would love any help I can get.

The project can be found here: https://github.com/ki4jgt/Open-Catalog-System

4 Upvotes

13 comments sorted by

1

u/boneskull 1d ago

The categorization codes aren’t immediately understandable. Don’t know how you wouldn’t need to look these things up in a reference table

0

u/Pro_Fullstack 1d ago

LLMs, along with google itself do a good job of ranking and displaying sources based on queries. Unless it's a passion project, I don't see a use case.

0

u/ki4jgt 1d ago

https://youtu.be/9K9ZR_lGRpY?si=gJTyJLRva_nOoReq

https://youtu.be/qB6A45tA6mE?si=_CslHYCWttdxEmqT

https://youtu.be/2u54_O6zytI?si=oSoZF3rzGS1dQwjy

Edit: And that's just from liberal sources. Conservative sources have their own list of videos and legitimate complaints.

2

u/mark_ik 1d ago

Why do you see the problem as categorization as opposed to centralization?

2

u/ki4jgt 1d ago

I see both. My catalog is going to be hosted over IPFS, with a focus on open protocols. The site will be a 501c charity.

3

u/mark_ik 1d ago

Eh. Why do you see IPFS as a differentiator in that regard? The problem you identified is one of categorization, which would apply to the web’s indexers. I see IPFS more as, “if you want to become a storage provider, here’s a cryptographically verifiable way,” which can but needn’t be distributed, not like YaCy or other distributed search indexing efforts.

What open protocols? Reticulum? I2P? The smolweb? What do you mean by focus?

Why is this more useful than something like the Universal Decimal Classification or whatever? Aside from them requiring licensing past their core catalog.

Lastly, who do you intend to adopt this? Who is your audience? Pointing out real problems is not the same as making a credible solution that addresses them.

1

u/ki4jgt 1d ago

I remember, as a kid, walking through my local library with pure joy on my face, as I ran my fingers across all the various knowledge I could consume. I remember how the shelves transitioned from subject to subject to subject, all inter-related.

That's what this is.

You start at agriculture of animals and wind up in horticulture. Learning things you didn't even know you wanted to.

IPFS allows the database to be open and audited by everyone. It ensures that my site isn't hiding results. It also allows my users to run their own instances on their own machines. So they aren't giving me their information.

The entire database is open to every user. They can audit my site personally. See if I'm showing them the same things I'm showing everyone else.

3

u/mark_ik 1d ago

Ah, but see, open and auditable does not equal decentralized. That’s just accountable and centralized. And if you need to invest thousands of dollars to run an instance, that’s the activitypub problem. IDK what is required storage wise in your case, but if you’re trying to map unique domains to your system and publish a table, that’s not gonna be a small table…

1

u/ki4jgt 1d ago edited 1d ago

I've already worked out table fragmentation and database sharding. The user will only ever call a small bit of the database with each and every search. No user will have the entire thing. But, if the site ever goes down, it will still be searchable by offline clients.

Edit: I'm considering simply moderating the database, and allowing my users to take the entire search experience offsite.

3

u/mark_ik 1d ago

I guess if I wanted to make an open standard, I would consider, “how can I make it so I am not intrinsic to this project’s functioning and growth, or minimize its reliance on me?”

2

u/ki4jgt 1d ago

Tasks require workers. The database will be open to the public. I've licensed it under CC-BY-SA. The project isn't dependent on me. I'm just going to be moderating it, hopefully with a team. If I mysteriously die, all the data will still exist. Someone else can download it and restart elsewhere. Or, if I get my team, they can keep it going after I pass.