HN user

tkyjonathan

201 karma
Posts30
Comments204
View on HN
shatteredsilicon.net 3mo ago

The AWS Lambda 'Kiss of Death'

tkyjonathan
4pts1
www.youtube.com 6y ago

Ultraviolet light to disinfect against Coronavirus

tkyjonathan
7pts1
news.ycombinator.com 6y ago

New Term: “Dark Cloud”

tkyjonathan
9pts7
github.com 7y ago

Primacy-Based Todo List

tkyjonathan
2pts0
www.youtube.com 7y ago

Jordan Peterson’s guide to leadership

tkyjonathan
7pts0
www.livekindly.co 7y ago

AI company in Chile creates vegan mayo

tkyjonathan
1pts0
www.youtube.com 7y ago

How SQL Databases Come Up with Algorithms That You Would Have Never Dreamed Of

tkyjonathan
2pts0
www.youtube.com 7y ago

Critical Computational Empowerment

tkyjonathan
1pts0
www.youtube.com 8y ago

Companies die when they run out of creative people (YouTube)

tkyjonathan
13pts5
medium.com 8y ago

How to Implement Technical Change in an Organisation

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

CST/Cogs Framework – IT Org Principles for Craftsmanship and Innovation

tkyjonathan
2pts0
www.jonathanlevin.co.uk 8y ago

MySQL Linux Tuning Checklist

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

A DBA Analyses 'The Phoenix Project'

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

Top 4 Reasons Companies Won't Fix Their Database Issues

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

Setting Up Databases in Your Development Environment

tkyjonathan
2pts0
www.theguardian.com 8y ago

Meat tax ‘inevitable’ to beat climate and health crises, says report

tkyjonathan
44pts78
www.jonathanlevin.co.uk 8y ago

Archiving for a Leaner Database

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

How to Not Be the One That Deploys That Slow Query to Production

tkyjonathan
1pts0
www.jonathanlevin.co.uk 8y ago

Top 3 Reasons Why SQL Is Faster Than Java

tkyjonathan
2pts0
www.jonathanlevin.co.uk 8y ago

What Is a Good Data Model

tkyjonathan
1pts0
ama.com.au 8y ago

Processed meats need a closer look

tkyjonathan
1pts0
help.gooddata.com 8y ago

Optimizing Data Models for Better Performance

tkyjonathan
4pts0
smalldatum.blogspot.com 9y ago

Sysbench for MySQL 5.0, 5.1, 5.5, 5.6, 5.7 and 8

tkyjonathan
66pts32
twindb.com 9y ago

RDS vs. Aurora vs. EC2 benchmark

tkyjonathan
2pts0
www.jonathanlevin.co.uk 9y ago

MariaDB's Columnar Store

tkyjonathan
4pts0
www.jonathanlevin.co.uk 9y ago

JSON and MySQL Stored Procedures

tkyjonathan
4pts0
www.jonathanlevin.co.uk 10y ago

How to Speed Up the Database Behind Your API

tkyjonathan
3pts0
github.com 10y ago

Collection of Clinical Studies about Vegan, Vegetarian and Meat-Eating Diets

tkyjonathan
2pts0
news.ycombinator.com 11y ago

Could we help Greece?

tkyjonathan
14pts14
latestvegannews.com 11y ago

Google Adding Plant-Based Options to Its Worldwide Food Program

tkyjonathan
1pts0

Almost none of the company leaders or even VCs fully understand what AI even is or does. They just like to hear thats its there.

If you don't have some AI in your company, you won't get investors.

"Normalization was built for a world with very different assumptions. In the data centers of the 1980s, storage was at a premium and compute was relatively cheap."

But forget to do normalisation and you will be paying 5 figures a month on your AWS RDS server.

"Storage is cheap as can be, while compute is at a premium."

This person fundamentally does not understand databases. Compute has almost nothing to do with the data layer - or at least, if your DB is maxing on CPU, then something is wrong like a missing index. And for storage, its not like you are just keeping old movies on your old hard disk - you are actively accessing that data.

It would be more correct to say: Disk storage is cheap, but SDRAM cache is x1000 more expensive.

The main issue with databases is IO and the more data you have to read, process and keep in cache, the slower your database becomes. Relational or non-relation still follows these rules of physics.

This is obvious to me. Since Hadoop came out, (a lot of) people have been giving up on even forming algorithms and just dumping data into machine learning and hoping for the best. I recall someone high up at Google complaining about it.

We need to get back to forming algorithms as well as concepts and first principles. We cannot and should not expect ML to brute force finding patterns and just sit back and relax.

Here is another prediction for you: we will not solve ray-tracing in games and movie CGI with more hardware. We will need some algorithm that gets us 80-90% of the way there in a smart way.

If its minimum of 70k, then its probably a good idea, because you would only hire if you absolutely have to.

37signals would be proud.

If you decentivize the most productive people, then you might hurt the rest of the company.

Batching is the multi-threadedness of databases.

Its also important to remember that in databases, you are more often optimising for IO usage than CPU.

Not sure which DB you are using, but you can load the csv file into the DB directly on a single thread using something like LOAD DATA INFILE.

If you have some good indexes and do some push-down work (give the database aggregation tasks to do instead of your python code), you should probably be more than fine.

For a 250Gb file.. should be ok.. maybe add some partitioning too.

Having Kids 7 years ago

Think about having kids as like starting your own start up which you want to invest in for years to come and will very likely IPO with huge investment one day.