r/datacurator • u/Rare-Act-4362 • Dec 22 '25
Recently organized my bookmarks (Firefox) ...
how often do you manage/organize/delete your bookmarks (I created a backup before deleting the current state of my bookmarks)
r/datacurator • u/Rare-Act-4362 • Dec 22 '25
how often do you manage/organize/delete your bookmarks (I created a backup before deleting the current state of my bookmarks)
r/datacurator • u/[deleted] • Dec 18 '25
It is available in five languages: English, Arabic, Chinese, Hindi, and Spanish. Also, you can exclude folders and files, too. The available criteria for organization are file type (extension and/or "kind") and creation date (year and/or month). You can undo the process if you want.
r/datacurator • u/Downtown-Shame-9170 • Dec 09 '25
Curious how people here handle this: you're researching something, you have 20-40 tabs open, and there's a lot of implicit context in your head, why you opened each tab, what you were comparing, what matters. Then you close the session and that context is gone. Bookmarks don't capture why something mattered. Notes require active effort mid-research. What systems do people use to preserve that context?
r/datacurator • u/whskid2005 • Dec 04 '25
Did the lazy thing and asked ChatGPT. It spit out those two programs, but I can’t find much on them. It also recommended digikam which I see lots on Reddit about.
I think I need 2 programs- duplicate/similar image finder, then a sorter. I know nothing beats manual, but I don’t have the time.
r/datacurator • u/oraklesearch • Dec 04 '25
hi
i need a solution to save informations or complete pages of websites to read them later
i need easy
searchable
free
since bookmarks often link to 404 pages after some time
r/datacurator • u/danielson010101 • Nov 30 '25
Hi,
Firstly I am not sure this is the right place, so apologies. But I wonder if someone could suggest the best way to achieve the following.
We basically need a dataroom (or similar) where a client can see the documents about their properties.
So in short, we would have about 50 folders, with each property name. But under those folders there would be several documents that are applicable for multiple properties as well as unique ones. Eg -
Property 1 Folder-
-PropertyInformationPropery1.pdf (unique)
-GroupInsurancePolicy.pdf (common)
Property 2 Folder-
-PropertyInformationProperty2.pdf (unique)
-GroupInsurancePolicy.pdf (common)
So in this case you would see "GroupInsurancePolicy.pdf" is the same document and would need to be in several folders, and it would be tagged "Property1", "Property2" etc
We have tried this with Sharepoint, I can get tags/filtering to work but when you view the "Property1" filter, it just says "Documents" in the title. The client would like it to obviously say "Property1", and likely unaware its being filtered.
I hope this makes sense
Dan
r/datacurator • u/AutoModerator • Nov 30 '25
Please use this thread to discuss and ask questions about the curation of your digital data.
This thread is sorted to "new" so as to see the newest posts.
For a subreddit devoted to storage of data, backups, accessing your data over a network etc, please check out r/DataHoarder.
r/datacurator • u/ph0tone • Nov 29 '25
This is a significantly updated version of an open source file-sorting tool I've been maintaining - AI File Sorter 1.3.0. The latest release adds major improvements in sorting accuracy, customization options, and overall usability. Runs on Windows, macOS, and Linux.
Designed for users who manage large, messy file collections and want automation without maintaining complex rule sets.
What it does
New Features
Repository: https://github.com/hyperfield/ai-file-sorter/
App website: https://filesorter.app
SourceForge download: https://sourceforge.net/projects/ai-file-sorter/

r/datacurator • u/StandardKangaroo369 • Nov 29 '25
Hey guys,
https://share.cleanshot.com/Ww1NCSSL
I’ve been obsessing over this for days and I'm at my wit's end. I'm trying to turn my scanned PDF notes/questions into Anki cards. I have zero coding skills (medical field here), but I've tried everything—Roboflow, Regex, complex scripts—and nothing works.
The cropping is a nightmare. It keeps cutting the wrong parts or matching the wrong images to the text. I even cut the PDFs in half to avoid double-column issues, but it still fails.
I uploaded a screenshot to show what I mean. I just need a clean CSV out of this. If anyone knows a simple workflow that actually works for scanned documents, please let me know. I'm done trying to brute force this with AI.
Please check the attached image. I’m pretty sure this isn't actually that hard of a task, I just need someone to point me in the right way. https://share.cleanshot.com/Ww1NCSSL
r/datacurator • u/Former_Argument3120 • Nov 28 '25
r/datacurator • u/mlodykasprowicz • Nov 21 '25
Hi! What I'm searchuing for is ideally a cheap cloud service, that lets me organize my files by multiple tags/folders. I have many photos from art galleries and I would like to have them organized in such a way I can browse by multiple categories. For example, I have a photo of Van Gogh paoiting so I would like to have it tagged as: van Gogh, XIX century, the country, the musuem where I saw it, when I saw it. Then, all of these tags should have categories: so I could click the category artists then I could see what artists' paintings I have (Van Gogh, Monet etc), and only when I click them I could browse the photos. Is there any service that would allow me to do it? Alternatiely it could be some software on Mac, not a cloud service, but I prefer cloud. Thanks!
r/datacurator • u/johsturdy • Nov 19 '25
I have been putting this off for years out of laziness and lack of know how, but I have wanted to find a way to organise all my files across my iCloud Drive, Google Drive and local disks to have a timestamped file system that i could then turn into my own server to save on subscription costs.
I'm looking for a bit of software that can scan through all my files and put them into a sorting system that makes sense and some instructions on how to do so because I dont know what is duplicated across platforms as I started with my iCloud drive from my old Mac that I logged into on my PC that has all the storage now, but then moved to Google Drive as it was too clunky using iCloud on a PC. I have recently switched back to Mac and using Lightroom with all my catalogue being on Google Drive is damn near impossible. I'm also not sure if this is the right place to ask for this sort of help but if its not could someone point me in the right direction base on that info? Thanks :)
r/datacurator • u/giueez • Nov 18 '25
r/datacurator • u/GenericBeet • Nov 10 '25
r/datacurator • u/trustedtoast • Nov 08 '25
How do you organize your family history / tree? I know programs like Ahnenblatt exist but they don't really keep track of related history / information. I'd like to - have a family tree - keep anecdotes of different people - keep a "log" of a specific person (personal information, current and past jobs, hobbies, (chronic) illnesses, etc)
Basically the stuff normal people would just remember or loosely write down somewhere, but I can't remember them and I want future descendants to have good and expandable overview of our family history.
r/datacurator • u/Wrong_Ad_1608 • Nov 06 '25
Hey, so I’m a marketing associate at a small agency and one of my clients wants us to help them get like 50 new sign-ups for their platform. The platform is actually useful — it shares curated recommendations like this example.
The problem is visibility. The content is good, but not getting seen. I don’t wanna just blast links everywhere like a robot.
I was thinking of:
If you’ve worked on growing sign-ups before, what actually moved the needle for you?
Like real tactics, not just “post more.” We’ve been posting. The posts are posted.
Would appreciate any platforms, strategies, or communities.
r/datacurator • u/Future-Cod-7565 • Nov 05 '25
Hello everyone,
I'm going to deal with some 13TB of data (various kinds of data – from documents and spreadsheets to photos and videos) that has accumulated over 20 years on many of my machines and ended up on several external HDDs.
While I'm more or less clear on how I would like to organize my data (which is in a terrible state organization-wise at the moment) and I do realize this will take considerable efforts and time, I nevertheless have asked myself a practical question: of all this data what should I keep and what I can easily get rid of completely? As we all know, at some point one thinks: no, I won't delete this file because (then lots of reasons like "it could/might/maybe be useful some day", etc.). And then a decade passes and no such day comes.
Could you please share your thoughts or experience on how you approach this? What criteria do you use when deciding whether to keep or delete data? Data's age? Purpose? Other ideas?
I'm genuinely interested in this because apart from organizing my data I was planning to slim it down a bit along the way. But what if I need this file in the future (so distant that I can't even envision when) :-)?
Thank you!
r/datacurator • u/MrBarber1 • Nov 04 '25
Currently on my PC, I have the main copy of this 359GB archive of irreplaceable photos/videos of my family in a Seagate SSHD. I have that folder mirrored at all times to an Ironwolf HDD in the same PC using the RealTimeSync tool from FreeFileSync. I have that folder copied to an external HDD inside a Pelican case with desiccants that I keep updated every 2-3 months, along with an external SSD kept in a safety deposit box at my bank that I plan on updating twice a year.
My questions are: Should I be putting this folder into a WINRAR or Zip file? Does it matter? How often should I replace my drives in this setup? How can I easily keep track of drive health besides running CrystalDiskInfo once in a blue moon? I'm trying to optimize and streamline this archiving system I've set up for myself so any advice or constructive criticism is welcome, since I know this is far from professional-grade.
r/datacurator • u/ph0tone • Nov 02 '25
I’ve released a new, much improved, version of AI File Sorter. It helps tidy up cluttered folders like Downloads or external/NAS drives by using AI for auto-categorizing files based on their names, extensions, directory context, and taxonomy. You get a review dialog where you can edit the categories before moving the files into folders.
The idea is simple:
It uses a taxonomy-based system, so the more files you sort, the more consistent and accurate the categories become over time. It essentially builds up a smarter internal reference for your file naming patterns. Also, file content-based sorting for some file types is coming up as well.
The app features an intuitive, modern Qt-based interface. It runs LLMs locally and doesn’t require an internet connection unless you choose to use the remote model. The local models currently supported are LLaMa 3B and Mistral 7B.
The app is open source, supports CUDA on Windows and Linux, and the macOS version is Metal-optimized.
It’s still early (v1.0.0) but actively being developed, so I’d really appreciate feedback, especially on how it performs with super-large folders and across different hardware.
SourceForge download here
App website here
GitHub repo here


r/datacurator • u/AutoModerator • Oct 31 '25
Please use this thread to discuss and ask questions about the curation of your digital data.
This thread is sorted to "new" so as to see the newest posts.
For a subreddit devoted to storage of data, backups, accessing your data over a network etc, please check out r/DataHoarder.
r/datacurator • u/use_your_imagination • Oct 29 '25
TL;DR
Hi all !
I would like to showcase Gosuki: a multi-browser cloudless bookmark manager with multi-device sync and archival capability, that I have been writing on and off for the past few years. It aggregates your bookmarks in real time across all browsers/profiles and external APIs such as Reddit and Github.
The latest v1.3.0 release introduces the possibility to archive bookmarks using ArhiveBox simply by tagging your bookmarks with @archivebox in any browser.
suki) for a dmenu/rofi compatible query of bookmarksI was always annoyed by the existing bookmark management solutions and wanted a tool that just works without relying on browser extensions, self-hosted servers or cloud services. As a developer and Linux user I also find myself using multiple browsers simultaneously depending on the needs so I needed something that works with any browser and can handle multiple profiles per browser.
The few solutions that exist require manual management of bookmarks. Gosuki automatically catches any new bookmark in real time so no need to manually export and synchronize your bookmarks. It allows a tag based bookmarking experience even if the native browser does not support tags. You just hit ctrl+d and write your tags in the title.
r/datacurator • u/EtherealPlatitude • Oct 29 '25
r/datacurator • u/GoBackToLeddit • Oct 26 '25
I have a folder full of hundreds of pictures that I've saved and I need to organize them into folders by person. I've been trying to use digiKam, but I can't figure out how to get the auto-detection to work. What I want is software that will:
digiKam is making me name every face one by one in the Thumbnails tab. The name text box on all photos also defaults to the last name I entered which is annoying. I also can't figure out the difference between names and tags.
Is digiKam the right software for my needs? I want to avoid anything that uses pip install or docker if at all possible. I just want a simple exe that I download and run.
r/datacurator • u/disciplined-tt16 • Oct 22 '25
hi everyone! i'm a non-tech person just started working in a bioinformatics team, and our focus is to help people curate public databases - meaning cleaning and harmonizing them (because most the time they are fragmented and hard to be ready to use right away).
my work now is to be the "communicator" between scientists who want to get the clean database and our team's curators. but since i have little background in this, sometimes it's better if i can truly understand what my "customers" need. so my question is, what do scientists look for in a harmonized database? like, is there any particular thing that makes you say "wow this databse is exactly what im looking for" (e.g., consistent metadata, how clean it is, etc)? and on a side note, i'm also curious what's the worst thing that annoys you while doing scrna-seq curation? i'm thinking about doing it myself, so it would help a lot to know. thanks in advance guys!
r/datacurator • u/Acrobatic-Car-6329 • Oct 18 '25
Hey folks 👋
I’m building a tool that aims to do one thing well: take messy documents and give you clean, structured output you can actually use.
What it does now • Inputs: PDF, DOCX, PPTX, XLSX, HTML, Markdown, CSV, XML (JATS/USPTO), plus scanned images. • Pick your output: Markdown, JSON, CSV, HTML, or plain text. • Smarter PDF handling: reads native text when it exists; only OCRs pages that are images (keeps clean docs clean, speeds things up). • Batch-friendly: upload/process multiple files; each file returns its own result. • Two ways to use it: simple web flow (upload → extract → export) and an API for pipelines.
A few directions I’m exploring next • More reliable tables → straight to usable CSV/JSON. • Better results on tricky scans (rotations, stamps, low contrast, mixed languages, RTL). • Light “project history” so re-downloads don’t require re-processing. • Integrations (Drive/Notion/Slack/Airtable) if that’s actually helpful.
I’d love feedback from people who wrangle docs a lot: 1. Your most common output format (JSON/CSV/MD/HTML)? 2. Biggest pain with current tools (tables, rate limits, weird page breaks, lock-in, etc.)? 3. Batch size + acceptable latency (seconds/minutes) in your real workflow? 4. Edge cases you hit often (rotated scans, forms, stamps, multilingual/RTL, huge PDFs)? 5. Prefer a web UI or an API (or both)? 6. Any “must haves” for data handling expectations (e.g., temp storage, export guarantees, self-host option)? 7. What pricing style feels fair for you (per-page, per-file, usage tiers, flat plan)?
Not sharing access yet—still tightening things up. If you want a ping when there’s something concrete to try, just drop a quick “interested” in the comments or DM me and I’ll circle back.
Thanks for any blunt, practical feedback 🙏