HN user

dbrereton

4,951 karma

Serial Side Project Maker

Posts111
Comments51
View on HN
lastmuseum.com 1mo ago

Show HN: The Last Museum – Semantic search across all museum art

dbrereton
3pts0
talesoftimesforgotten.com 2mo ago

Did Ancient Civilizations Have Organized Crime?

dbrereton
3pts1
calnewport.com 2mo ago

Easy Is Overrated

dbrereton
4pts1
unchartedterritories.tomaspueyo.com 3y ago

How to Create a Masterpiece

dbrereton
2pts0
podscript.ai 3y ago

PodScript

dbrereton
1pts0
erikhoel.substack.com 3y ago

The White House agrees you have a small brain

dbrereton
4pts1
danieljeffries.substack.com 3y ago

Rise of the AI Doomsday Cult

dbrereton
3pts0
yyyyyyy.info 3y ago

Yyyyyyy.info

dbrereton
2pts2
rmitz.org 3y ago

The FreeBSD Daemon

dbrereton
2pts1
itwasntai.com 3y ago

It Wasn't AI

dbrereton
131pts129
dkb.blog 3y ago

AI Will Make the Internet Beautiful

dbrereton
3pts0
blog.samaltman.com 3y ago

The days are long but the decades are short (2015)

dbrereton
115pts83
dkb.blog 3y ago

The War Against AI Art

dbrereton
2pts0
pedestrianobservations.com 3y ago

New York Can’t Build, LaGuardia Rail Edition

dbrereton
1pts0
dkb.blog 3y ago

ChatGPT's Chess Elo is 1400

dbrereton
212pts345
www.lesswrong.com 3y ago

Simulators

dbrereton
1pts0
dkb.blog 3y ago

TikTok: The Search Engine for Experiences

dbrereton
1pts0
pmarca.substack.com 3y ago

Welcome

dbrereton
4pts0
www.axios.com 3y ago

AI and ChatGPT are selling cars in the metaverse

dbrereton
1pts1
www.strangeloopcanon.com 3y ago

Don't Panic: Against the Butlerian Jihad on AI

dbrereton
2pts1
walkingtheworld.substack.com 3y ago

Walking South LA

dbrereton
2pts0
pedestrianobservations.com 3y ago

Cost and Quality

dbrereton
2pts0
dkb.blog 3y ago

ChatGPT Fails the Coding Interview

dbrereton
3pts0
talesoftimesforgotten.com 3y ago

ChatGPT Is Impressive for a Bot, but Not for a Human

dbrereton
1pts1
bitsofwonder.substack.com 3y ago

Death by Over-Recommendation

dbrereton
1pts0
en.wikipedia.org 3y ago

Tay (Bot)

dbrereton
1pts1
www.maa.org 3y ago

A Mathematician’s Lament [pdf]

dbrereton
1pts0
dkb.blog 3y ago

Bing AI can't be trusted

dbrereton
1072pts601
talesoftimesforgotten.com 3y ago

Aristotle Was Not Wrong about Everything

dbrereton
2pts0
noahpinion.substack.com 3y ago

Vertical Communities

dbrereton
1pts0
It Wasn't AI 3 years ago

Seems like the site might have crashed, here's an archive link: https://web.archive.org/web/20230515030802/https://itwasntai....

But tl;dr many students have been accused of using AI by teachers who think that AI detection software works, when it really doesn't. So the goal of this site is to communicate to teachers that AI detection software isn't reliable.

I originally discovered this in a reddit comment which you can see here: https://www.reddit.com/r/ChatGPT/comments/13hi5y6/comment/jk...

I've been working on a manual version of this for many months now [0]. I read books of historical figures, then create fictional but accurate conversations with them based on the books. And I always include full citations so people know it's accurate.

I think releasing a purely AI version of this is not feasible at the moment, because of the obvious issue of hallucinations. It's not an educational product if it's giving people false information. In fact, it's actively harmful. And it's sad that people are trying to cash in on the AI hype without any regard for the accuracy of their content.

[0] https://dkbshow.substack.com/

Yeah I think the bigger issue here is not highlighting the matched text, which makes it look like it's doing something totally random.

There is some typo tolerance, so it's hard to tell exactly what is matching on those pages.

I'm working on adding query operators so you can specify things like exact matching, and in general working on improving the results.

This type of feedback is very helpful, so thank you!

My understanding is that the early days of blogging were mostly people sharing personal things, and now that has moved to social media.

A lot of the online writing happening now is essays on some topic, or people trying to share notes on things they learn. And I think this type of writing is more conducive to a project like this.

The simple and unfortunate explanation is that the index is just not that big right now (only 900 blogs).

Working on increasing it significantly, but will take some time. Try it again in a month and you may find it more useful. Right now it's mostly filled with tech, business, and politics.

Only 900 blogs are currently indexed, and they're mostly tech, business or politics, so I wouldn't be surprised if none of them have written about the will smith situation.

I am working on increasing the amount of blogs significantly, but please bear with my modest index in the meantime.

Thank you!

changing a search phrase or word and doing a new search, I notice the results do change, but there's no way to know if it really happened. Changing a Google search, the whole page flashes empty, that way I see/sense there's something new. In your case, a change is subtle, very subtle, too subtle. In one instance I had to look carefully to see the change in results.

That is a good point. I did try to remove any loading indicators because I thought it would be smoother, but maybe it's a bit too smooth for people to realize their search went through. Will think more on how to fix this.

The value of a content shall be in the content itself; popularity is only a flawed measurement.

Popularity is certainly a flawed measurement, but it's hard to come up with a scalable way to determine quality that isn't flawed in some way.

Instead of being flawed by encouraging people to get tons of backlinks, this is flawed by encouraging people to do stuff that gets lots of upvotes.

Very open to more ideas on how to measure quality.

This is not a business, and I would like to open source it, but it would probably be better for everyone if I wait until I clean up my garbage code, which will take some time.

Right now the "random interesting posts" are a random selection from the top 1000 of all time.

However, if you want fresher content, you can use the date range selector and set it to "Past Week" for the best posts of the week.

Yeah tagging is definitely a weak point as the current focus is on search. I just removed the tag requirement.

If you're submitting a blog and we don't already have a tag that fits, you can add it to the "Notes" section.

Tbh I'll probably use the random bit more than search

That's interesting to hear, and fits well with the goals of the site. I want it to be more of a "discovery engine" than a "search engine". Search is one path to discovery, random posts are another, there are probably more.

One thing I'm thinking of adding is the ability to easily see the blog posts that any given post links to. If you see an interesting post, you could pull up everything that may be related.

definitely going to keep checking back to pad my RSS feeds with interesting content.

Sadly not every blog has RSS, and many RSS feeds are incomplete. Another thing I would like to build is auto-generated RSS feeds for all blogs, which would also make it easy for people to programmatically parse any blog and do interesting things.

This is definitely very cool as I've been looking for something like this since technorati (which was originally a blog search engine).

Technorati was one of the inspirations here so that's great to hear.

Would love to hear details about how you created the database, the infrastructure, etc if it's not a trade secret. Kudos on the launch!

Sure, it's actually fairly simple! The search backend itself is running on Typesense [0], which was very quick and easy to setup.

Due to the way ranking is calculated, I can actually avoid doing any real web crawling (though, I may add that in soon to help increase the index size). Ranking is based on submission to online communities, so all I really need is those submissions.

Using the Reddit, HN and Twitter APIs, I search for any submissions related to any blogs in the database, then those submissions give me the post URLs.

Once I have the post URLs, I just need to request those specific URLs to get the post data.

Then there's scripts for things like content extraction, inflation calculation, currency conversion etc.

All of those scripts are in python.

The frontend is a simple React app built with Next. All pages are statically generated.

Let me know if there's any more questions!

[0] https://typesense.org/

Ah yeah, there's probably not any blogs that talk about video editing. The index is fairly small right now and does better for tech/business/politics queries at the moment. Will work on increasing the index.

But I don't feel like manual curation by one person is easily compatible with search engine.

I see the manual curation as more of a temporary measure in the beginning. There are various ways blog detection can be automated and scaled, but manual curation for now gives me a better understanding of the data, and ensures I don't run into random edge cases.

Because after trying a few search "getting a job in vc", "best computer chair", "learning erlang" I'm not confident this answer better results than Google.

Right now it's more useful for very broad queries like "inflation" or "covid". The index is pretty small at the moment, but the more posts that get added to the index, the more specific queries we'll be able to find good results for.

You've got a content size problem as you are manually curating, and this will lead to people not use your search as a default, and probably not use it as a search engine, but instead as a discovery system.

That's actually what I want! This is not a search engine to replace Google, it's a discovery tool for blog posts.

Thanks for all the feedback here, and will definitely check out the newsletter.

I totally agree with you on this. Being able to follow the links easily in either direction would make it easier to fall into interesting rabbit holes.

Edit: hmmm... though looking further, maybe this goes against your MarketRank philosophy.

It doesn't at all, but curious as to why you'd think that. We're not talking about using backlinks to rank pages after all, just as a discovery tool, which I think is great.

Lots of great new search engines popping up that search the rest of the web that Google tends to ignore.

Other ones worth checking out include:

- https://search.marginalia.nu/ (A non-commercial search engine)

- https://wiby.me/ (Tends to have those really weird and cool indie sites)

- https://searchmysite.net/ (An index of personal websites)

- https://indieweb-search.jamesg.blog/ (Search IndieWeb websites)

- https://millionshort.com/ (Ignore the first million results from Google)

It’s been hard for me to find time to read books, and there’s a ton of things I want to read, so I built Blog Books to solve this.

Blog Books makes it easier to read more books by bringing the books directly to your inbox that you already check every day.

You choose how much you want to read (e.g. 500 words), and how often (e.g. weekly), and the book contents will be sent to your email inbox or RSS reader on that schedule.

It’s mostly limited to books in the public domain right now, though there are a few recent ones as well. Hoping to get a lot more books in this format in the future.

What have you tried so far, and what’s the usage like? If lots of people are using it and genuinely find it valuable then maybe some portion of those would be willing to pay for it. Otherwise some type of sponsorship / non intrusive ads might work though it looks like you want to avoid ads.

Check out listen notes [0] for some inspiration. It’s a podcast search engine that is profitable, and the founder has some good write ups on the blog.

0. https://www.listennotes.com/

RIP Google Reader 5 years ago

As someone who was born too late for Google Reader, I genuinely don't understand why people bring it up every 5 minutes.

There are 1000 feed reader apps that exist right now, some of which have the branding of "it's just like Google Reader", so what am I missing here?