HN user

george3d6

543 karma

Blog: https://cerebralab.com

Github: https://github.com/George3d6

I work on Mindsdb (Opensource auto ML): https://github.com/mindsdb/mindsdb

Posts56
Comments36
View on HN
news.ycombinator.com 10mo ago

WTF is up with everything using Android?

george3d6
14pts28
epistemink.substack.com 2y ago

Increasing IQ Is Trivial

george3d6
1pts9
www.epistem.ink 3y ago

In Defense of Making Money

george3d6
1pts0
www.epistem.ink 4y ago

Eating Boogers

george3d6
2pts0
www.epistem.ink 4y ago

Differences in Civic and Legal Attitudes Towards Drugs

george3d6
1pts0
www.epistem.ink 4y ago

Surviving Automation In The 21st Century – Part 1

george3d6
1pts0
www.epistem.ink 4y ago

Autopilot Ethics and the Illusory Self

george3d6
1pts0
www.epistem.ink 4y ago

The Limits of Medicine – Part 1 – Small Molecules

george3d6
1pts0
cerebralab.com 5y ago

Revelation and Mathematics

george3d6
1pts0
cerebralab.com 5y ago

Google Ain't Worth a Dime

george3d6
1pts0
cerebralab.com 5y ago

The Web Development Pattern

george3d6
1pts0
cerebralab.com 5y ago

I don't want to listen, because I will believe you

george3d6
1pts0
cerebralab.com 5y ago

Machine learning could be fundamentally unexplainable

george3d6
2pts0
blog.cerebralab.com 5y ago

PSA: Beware of Temporary Remote Work

george3d6
2pts0
blog.cerebralab.com 5y ago

Determining Determinism Is Indeterminable

george3d6
1pts0
blog.cerebralab.com 5y ago

Keeping Encryption Elitist

george3d6
2pts4
blog.cerebralab.com 5y ago

Costs and Benefits of Metaphysics

george3d6
1pts0
cerebralab.com 5y ago

Is a billion dollars worth of server lying on the ground?

george3d6
330pts331
blog.cerebralab.com 5y ago

Exploring the human condition via 3d6 is a row

george3d6
2pts1
blog.cerebralab.com 5y ago

Dual licensing GPL for fame and profit

george3d6
114pts112
blog.cerebralab.com 5y ago

Exceptions as Control Flow

george3d6
3pts0
blog.cerebralab.com 5y ago

Code as Pointless as Tea

george3d6
1pts0
blog.cerebralab.com 5y ago

Musings on the Impossibility of Testing

george3d6
27pts50
cerebralab.com 5y ago

Longevity Interventions When Young

george3d6
2pts0
cerebralab.com 6y ago

Science Eats Its Young

george3d6
3pts0
cerebralab.com 6y ago

Confidence in Machine Learning

george3d6
1pts0
blog.cerebralab.com 6y ago

Causality and Its Harms

george3d6
1pts0
blog.cerebralab.com 6y ago

Training our humans on the wrong dataset

george3d6
1pts0
blog.cerebralab.com 6y ago

Is a trillion dollars’ worth of programming lying on the ground?

george3d6
299pts334
blog.cerebralab.com 6y ago

Abstraction isn't wrong, it's just bad

george3d6
1pts0

Expand? The design constraints seem very dissimilar: - weight is important, form factor isn't - any app is either doing 3d rendering or integrated into a windowing system ala a trad desktop - they are not constantly-on, they are used for long periods with breaks - they are limited by processing power not by UX - the peripherals are very complex - to the extent of complex requiring access to hand and/or controller movement at a very fine level

If anything they are similar to laptops, but overall they are their own device class.

If "similar to mobile phones" means "similar chips and batteries"... sure, but so are laptops build after 2020, besides that I don't quite understand the comparison.

You’re not a mega company Right, I'm asking why Steam, HTC and Meta did this, which are in the top <x> tech companies by market cap and spend billions on this stuff. Totally get why a small VR device dev would not do this.

OS, drivers and libraries for (3D) graphics up and running Right, except that I am not saying one should build a kernel, and webgl, cuda, etc are kernel dependant not OS dependant

assurance that you still can get hardware to deploy your stuff on X years from now I am unaware of android-specific hardware features to date (linux specific, maybe, but not android) -- you can literally add any bootloader and load a linux kernel and get drivers running on most phones even, w/o android, and where "android" is necessary it's simply an artifact of the company maintaining the phone having build the drivers into their distro and not released them separately.

You also don’t want to order a million units up front. Again, seems unrelated, I'm talking about building the VR hardware i.e. not order off the shelf and white label (unless your claim here is that meta, htc, steam etc are doing white labeling, which doesn't seem to be the case)

it's more like, in a world with thousands of oncologists that treat pancreatitis cancer, saying: You should go to one, I'm not going to recommended mine in particular, but I went to one and she cured my cancer, so maybe you should give it a shot.

Did add a quick explanation about methods at the end ... I forgot that when I say "you can read the science around this topic to see how" and provide some links to examples of papers, people don't actually do it.

There are at least 6 cloud providers I can name that I've used which run their own data centers with capabilities similar to AWSs core products (ec2, route53, s3, cloud watch, rdb)

Ovh, scaleway, online.net, azure, gcp, aws

That's one's I've used in production, I've heard of a dozen more including big names like HP and IBM, I assume they can match aws for the most part.

...

That being said I agree multi tenant is the way to go for reliability. But I was pointing out that in this case even the simple solution of multi region on one provider was not implemented by those affected.

...

As for running your own data center as a small company. I have done it, buying components building servers and all.

Expenses and ISP issues aside, I can't imagine using in house without at least a few outages a year for anywhere near the price of hiring a DevOps person to build a MT solution for you.

If you think you can you've either never tried doing it OR you are being severely underpaid for your job.

Competent teams to build and run reliable in house infrastructure exist, and they can get you SLA similar to multi region AWS or GC (aka 100% over the last 5 years)... But the price tag has 7 to 8 figures in it.

This seems like an insane stance to have, it's like saying businesses should ship their own stock, using their own drivers, and their in-house made cars and planes and in-house trained pilots.

Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.

Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.

AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.

And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.

edit: I will add that AWS specifically is meh and I wouldn't use it myself, there's better IASS. But it's insanity to even imagine self-hosted is more reliable than using even the shittiest of IASS providers.

The market needs to open up for companies selling NFTs representing IpV4 address blocks. Not to provide any right to it's usage or anything, just for you know, the bragging rights.

You may not be the target audience though, the target audience are probably people that aren't yet "bough into" the google ecosystem. So they are unaware of this axing policy. Anyone logging into chromium, using gmail and google calendar and an android device is already "theirs".

What they want is to increase the share of people that are in the MS ecosystem (and presumably also OSX ecosystem, but that's mostly a status signaling thing, so the strategy there is probably different).

We are honestly very happy to hear about every usecase people have for this right now. It's users that help us shape the product to a large extent, at the end of the day.

We can chat more about it here or feel free to pick one of the many contact routes scattered around the thread :)

What's the name of the product? What does it do?

As far as I can tell MADlib is not automl, it just provides various statistic analysis and "classical" ml algorithms as function/macros in various database, and integration seems to be quite different from the way we do it (and I'd say more complex for the user, but maybe that's just my bias talking).

So I don't think there's a lot of overlap there. But if you think otherwise and work on it or know someone that does, I'd be quite excited to have a chat, just to share experiences and tips if nothing else.

If you're interested in participating in a "beta" release of that, we have a newsletter: https://mindsdb.com/newsletter/ where I'm 99% sure it will be announced, and in case you don't get a code ping me and I'll send you one :)

But the timeline for when to release this is still inexact, since ideally I want the "beta" to be as stable as possible, as to not have to get people to migrate to a different version later.

Also, if you have the time to share your usecase in like 3-4 paragraphs please do, either here or email me, because we design this with users in mind at the end of the day.

So I assume that you are doing hyperparameter search? Can you share what optimization method you are using for search (e.g. random, gp )?

Short answer is optuna and ax but only sometimes.

Long answer lead me down a rabbit whole and it's 10k+ words and a few experiments deep. If you're interested in this are specifically ping me, but I've got nothing concrete, however I like discussing it. A recent paper I saw that somewhat echos my thoughts is: https://arxiv.org/pdf/2102.03034.pdf | but some bits feel either over my head and/or overly pedantic and/or overly formal | and I'm not sure I agree with the conclusion | and loads of it is irrelevant. But if the problem interests you I'd suggest giving it some time, with those disclaimers in mind

Also, is the search can be distributed in parallel to multi node ?

Theoretically yes, practically it's still WIP to get this to work, but the architecture we have right now is very much conceived with massive distribution in mind (see our docs for more details on that).

And, if mindsdb is not part of the db, what happen if minddb fail ?

The select query you use to make a prediction returns an error, essentially. Assuming you mean "what happens if it crashes or if the model you are using crashes?".

e.g:

psql> SELECT diagnostic FROM mindsdb.flu_detector WHERE headache=true AND temperature=37.5 AND cough='mild';

psql> Error: External table returned error: "Segfault"

OR

psql> SELECT diagnostic FROM mindsdb.flu_detector WHERE headache=true AND temperature=37.5 AND coughsfsagsa='mild';

psql> Error: External table returned error: Input column `coughsfsagsa` doesn't exist

(or something like that)

Also, do you support automatic retraining?

Not at the moment, but we're going to add it very soon, with the first implementation allowing retraining with a certain user-set frequency (e.g. once every 2 hours).

Which will allow the model to be always fresh as new data comes in (assuming there's no time limit on the query)

Looks quite interesting, already pinned this in the relevant slack channel :)

To be honest I'm rather happy with how the internal benchmark suite is turning out, but to some extent you are inviting bias by creating them yourself. On top of that, it doesn't hurt to have more benchmarks.

At the end of the day it's a combination of: * How much work is it to integrate (easy to measure) * How visible is it, i.e if we actually find something interesting will be visible and legible to others (ify to mesure, citations, stars, etc are some invitation) * How useful it is to "improve" the library (hard to measure, and what we aim to be good at is a moving target)

So realistically that's the equation I have to judge in terms of adding a new benchmarks suite, and it's very annoying because you'll note the most important things are the hardest to measure.

Would you want people to integrate with this now or would you rather wait a few weeks/months/years until it matures more? If the former, can you give a few details regrading where to start (README is fairly barren), if the later please ping me (george.hosu@mindsdb.com) when you think it could be ready to try.

Anyway, any open benchmark library is a step in the right direction, thanks for working on this :)

From the user perspective it's inside the database, you can run mindsdb in the backgrond,connect it to the database once, and then do everything from within the database (i.e. connecting with a sql client to your database server and issuing commands the same way you would query "normal" tables).

From a technical perspective it's a separate server that communicates with the database through various mechanisms (e.g. federate engine) but it's no different from e.g. multiple instance of mariadb being abstracted by a galera cluster into something that behaves like a single database from a client perspective.

I would love to support Scylla, I ** love that database, those guys are magicians. And I assume in supporting that we'd also offer de-facto support for Cassandra.

I don't think either Scylla or dynamo are on the roadmap now, but if you want them feel free to create an issue asking for them: https://github.com/mindsdb/mindsdb

It should be noted that there's two level of support:

1. As a source of data (easy to implement) 2. Being able to publish models into the database (a bit harder)

If you work with those and are interested in doing ML from the database please get in touch, ideally via github, but you can also use the contact form (https://mindsdb.com/contact-us/) or email one of us directly. The best case scenario for us is that when we do one of these integrations we have an actual user in mind, and we're open to "first users" for any database where we can find a reasonable way of integrating.

I haven't looked into it myself, but I'll try to understand what they do better, thanks for letting us know.

I will say that:

1. It's not open source, so hard for us to compare other than running black-box experiments.

2. Oracle, so presumably that comes with all the Oracle-ecosystem buy-ins that implies, which might not be ideal for many people.

As a purely personal opinion:

I guess it's good to know that other people are thinking in the same direction as us, but at the same time I personally would like for widely-used ML libraries to be open-source. If these models are going to be used as generator of important decision making algorithms, ideally both the model and the algorithm should be open source. The later is up to whoever is building the algorithm, but I think if we can get the zeitgeist to move towards the later being open source as the norm that can alleviate a lot of potential harm and has little downside.

I.e. Do you feel comfortable with the NHS off-sourcing important decision making to algorithms that are proprietary black boxes? Considering that it's funded by the tax paying public and it's supposed to service that public.

"Secret" laws used to be a norm in the past e.g. in large civilziations like the Roman empire, where the norm evolved to be that only "schooled" men could understand the law due to complexity, or in most of medieval Europe where the bible was foundational for morality but closed off to a small subset of the population that knew Greek or Latin and could get their hands on it. But in general that seems to have caused more harm than good.

It seems reasonable to ask that, if algorithms are going to be used by governments in decision making, those should be entirely open. Ideally the ones used by corporations should be open to whatever degree is possible, to avoid run-off harm from buggy or unaligned systems.

Cheers :)

It's actually quite nice for me to hear that people we didn't hear from yet are finding it useful. Since it's a library it's hard to actually figure out how many people are really using it successfully and what for.

If you don't mind sharing more details please do, either through our usual channels (https://mindsdb.com/contact-us/) or just send me an email (george.hosu@mindsdb.com). Figuring out how people use it and what issues they encountered has been immensely helpful to me.

Regrading benchmarks, we have three main dataset collections we focus on currently:

1. Datasets from customers, but obviously those can’t be made public.

2. The OpenML benchmark, which is fairly limited because it’s mainly binary categories, but which is good because it’s a 3rd party, so unbiased. We have some intermediary results here (https://docs.google.com/spreadsheets/d/1oAgzzDyBqgmSNC6g9CFO...) , they are middle-of-the-road. However I think the benchmark is pretty limited, i.e. it doesn’t cover most of the kinds of inputs and almost none of the output we support

3. An internal benchmark suite which currently has 59 datasets, mainly focused around classification and regression tasks with many inputs, timeseries problems and text. Some part of it is public but opening that up is a bit difficult due to licensing issues. I’m hoping that in the next year it will grow and 90%+ of it can be made public. We benchmarkagainst older versions of mindsdb, against hand made models we try to adapt to the task, against the state of the art accuracy for the dataset (if we can find it) and a few other auto ML frameworks (well, 1, but I hope to extend that list) [see this repo for the ones we made public: https://github.com/mindsdb/benchmarks, but I'm afraid it's a bit outdated]

That being said benchmarking for us is still WIP, since as far as I can tell nobody is trying to build open source models that are as broad as what we're currently doing (for better or worst), and the closed source services offered by various IaaS providers don't really come with public benchmark results outside of marketing.

I don't think you're getting my point, consider reading again. There's a fundamental limit you will hit here, you can't just "make it slower" or "make it worst" to lower that limit.

Because "refine" really means "melt at temperatures ranging from 500 to 4000 degrees celsisu and then extract via various mechanical processes and/or using various reactants which often require gigantic plants to produce and are highly unstable".

Basically all of the history of science until 200 years ago was figuring out "mine and extract" and the course of civilization is very much linked with the price & quality of metal structures they could produce. But it's gotten so good we take it for granted.

However that is because of gigantic plants situated in specific areas where energy is cheap that do this thing at an amazing scale.

Aluminum is cheap as chips, except that it used to be more expensive than platinum (and at a much higher impurity ratio than the stuff we use for baking or for making cheap cases).

Heck, gold is "a thing" because we could purify and mold it without bringing it to a melting point and it was, for a very long time, the only metal available to us to do anything with, way before the bronze age.

And the problem with the refinement process is that you can't really "be smart" about it, reaching very high temperatures is one of those things you can't really scale down in an efficient way. You'd have to propel 100,000 tons of factory to mars in order to efficiently refine anything remotely close to the metals we had access to 100 years ago.

Which is not to touch on the mining bit, that is in itself very complicated (see how slowly and shallowly rovers are currently able to drill).

Are there workarounds for this? Maybe, I don't think anyone knows them though, they are not the kind of thing that's within easy reach. Maybe if we happen to stumble upon large reserves of bismuth or lead or gallium or mercury close to the surface of Mars, and build a whole branch of engineering around using those to build machinery... ? But my limited knowledge of geophysics and geology tells me that finding those in large amounts is very unlikely.

For reference, if you take an oven, that can reach, say, 450 degrees celsius (home) and up to 700 (industrial). Those aren't enough to refine any "useful" metal (e.g. iron) and building them requires materials that were produced at 1500+ degrees.

IANAChemist/IANAMaterialScientist/IANABlacksmith though, so take with a spoon of salt.

These sort of naturalist arguments seem to make little sense when thought about more clearly.

Of course that:

a) Destroying the environment is probably more destabilizing than keeping it intact... that's almost always the case with conservation vs change, conservation is the safe move

b) Destroying all insects (i.e a whopping 2/5th of all multi-celular bio diversity) would result in catastrophic damage... destroying 2/5th of all biodiversity would result in catastrophic damage.

The onus on any given insect conservation advocate is to prove:

1. Insects are more critical to the environment than other things we put more money and time into preserving (e.g Australia should spend more on preserving black widows and less on preserving koalas)

2. A very specific insect is critical and thus should be preserved at all cost, even if it means putting more money into the nature-preservation system

3. The bio-diversity of insects as a whole is dropping at an alarming rate compared to other species and this is harmful and unprecedented (biodiversity boom and bust cycles are the norm, usually)

I'm not saying the above 3 points can't be made.

But this article doesn't make them.

It sounds just like someone signaling "Hey, I'm a hippie ecologist, I love nature, even insects that sting, I'm that much of a group member, love me" rather than providing any new information or concrete policy proposal or even a meta-level argument for getting new information or creating new policy proposals.

The general idea of "we shouldn't destroy" nature has been around since, I assume, at least 5000 years ago or so, when bronze-age humans burnt down most forests in Europe and then realized this was bad and proceeded to replant and protect, or even venerate, forested areas. It's old new, nobody disagrees we should protect nature, it's just a typical tragedy of the commons problem, and one that hasn't been handled that badly thus far.

you don’t really do this if you care about the freedoms that GPL is meant to preserve

I mean, if this was the case, why provide the source at all ? Or why make it free to use forever ? Why not release it under a license that states usage is only permitted as long as the author allows it, i.e. for a limited trial period ?

My point here is that you can be pragmatic and say:

If some people want to build a free world, fine, I agree with that idea in principle and I will provide my work to them for free.

On the other hand, most people want a paid world, which is also fine, I might as well provide my service to them as well and benefit from it.

You're not doing as much as releasing stuff under GPL-only would, in that you're giving people that pay you an out, but it seems to be better than nothing. Plus, it's a more positive approach than GPL, in that it's not actively "hurting" people that don't want to join the open source community (by not providing them any option to use the product), it's simply giving an advantage to the open source users.

How does dual licensing work with 'downstream contributions'? If some user finds a bug in your GPL code, fixes it, and pushes that code to your repo?

You don't own the copyright to that bugfix right? So how can you re-license it under a commercial license?

The way we do it in my project, and the standard practice for Apache (where we copied it from), is tohave a CLA that gives full rights to the original owner for any patches people want to PR.

What about meaningful improvements instead of bug-fixes? I can see a bugfix being trivial enough. But if someone works hard to improve performance, and then some other company starts selling that work without compensation for the original author?

This is a bigger problem, but in practice I assume it wouldn't happen because if someone were to actually put in weeks or months of work into significantly improving the project, why wouldn't you just hire them or pay them ? After all, the whole assumption here is that this is a model for a for-profit endevor.

Yes, or rather, run your whole program with specific inputs then (e.g. by using a debugger) capture as much of it's state as you can and compare that with the state from the next run of those same inputs.

E.g. given a program with variables

a1 = array a2 = int a3 = string

Running might get the state flow:

a1 = []

a2 = 10

a3 = 'abc'

a3 = 'dd1'

a2 = 200

a1 = [200,'dd1']

Then, if a change is made, the programmer can say "I expect this change to only affect the state of a2" and if you get the state flow:

a1 = []

a2 = 11

a3 = 'abc'

a3 = 'dd1'

a2 = 500

a1 = [200,'dd1']

Then the test passes

If you get the state flow:

a1 = []

a2 = 11

a3 = 'abc'

a3 = 'dd1'

a2 = 500

a1 = [500,'dd1']

The test fails with error "The change you expected to only affect a2 also had an effect upon a1"

But again, the actual state you capture here and the way you check for equality (e.g. to you check for equality among all the states during the whole runtime, or just in the final state of the program ? If the former how is state-change ordering handled ? If the later, what about relevant state that is out of scope by the time the program finishes running ?)

What are your thoughts on automated/unit tests used to guarantee expectations do not change in future revisions (a regression suite) and also used to present example inputs usages of the API that was designed?

To be honest, this is one idea I was toying around with. As in, have a test suite that generates a "state" file for your program and then, in any given patch, list the items in the state that you'd expect to change.

That way, one can basically catch a lot of "bugs" which often boil down to "this change I made here to affect X is also unexpectedly affecting Y".

My main problem with this are:

1. I see nobody doing this, and I'm not sure why. 2. The tooling for this doesn't really seem to exist, and I'm not sure how easy one could bring it into existence and/or if a generic version of it could be written. 3. This breaks down with a lot of software where a tiny change can affect, well, everything (e.g. modifying a random seed or changing an error-checking constant)

I'm not trying to achieve perfect, rather just be able to answer a question like: What makes a test obviously bad or obviously good ? I needn't have a system that can help me generate perfect tests, or tell apart almost equally good implementations, but rather one where I can at least point to a test that is obviously bad and say "this is why" rahter than just using my intuition.