Cool to see principles behind this, although I think it’s definitely geared towards the consumer space. Shameless self plug, but related: we’re doing this for industrial assets/industrial data currently (www.sentineldevices.com), where the entire training, analysis and decision-making process happens on customer equipment. We don’t even have any servers they can send data to, our model is explicitly geared on everything happening on-device (so the network principle the article discussed I found really interesting). This is to support use cases in SCADA/industrial automation where you just can’t bring data to the outside world. There’s imo a huge customer base and set of use cases that are just casually ignored by data/AI companies because actually providing a service where the customer/user is is too hard, and they’d prefer to have the data come to them while keeping vendor lock-in. The funny part is, in discussions with customers we actually have to lean in and be very clear on “no this is local, there’s no external connectivity” piece, because they really don’t hear that anywhere and sometimes we have to walk them through it step by step to help them understand that everything is happening locally. It also tends to break the brains of software vendors. I hope local-first software starts taking hold more in the consumer space so we can see people start getting used to it in the industrial space.
HN user
montereynack
Gonna throw in my hat and say that if you’re working on industrial applications (like energy or manufacturing) give us a holler at www.sentineldevices.com! Plug-and-play time series monitoring for industrial applications is exactly what we do.
Gonna throw in my hat here, time series anomaly detection for industrial machinery is the problem my startup is working on! Specifically we’re making it work offline-by-default (we integrate the AI with the equipment, and don’t send data to any third party servers - even ours) because we feel there’s a ton of customer opportunities that get left in the dust because they can’t be online. If you or someone you know is looking for a monitoring solution for industrial machinery, or are passionate about security-conscious industrial software (we also are developing a data historian) let’s talk! www.sentineldevices.com
We’ve finally managed to give our AI models existential dread, imposter syndrome and stress-driven personality quirks. The Singularity truly is here. Look on our works, ye Mighty, and despair!
Seconding the other question, would be curious to know
‘"An ML engineer at Slack says they don’t use messages to train LLM models," Orosz wrote. "My response is that the current terms allow them to do so. I’ll believe this is the policy when it’s in the policy. A blog post is not the privacy policy: every serious company knows this."’
I really wish this advice was followed to the letter. I’m sick and tired of trying to read into policies on AI training (or AI anything these days) that is pure blog posts on how the service works and what the data protections are. Even from Microsoft no less! Even their “documentation” has a bad habit of referring to half-decisions alluded to in a blog post, saying this is how it works, just trust us, interpret this policy vaguely because we interpret it vaguely. None of the cloud vendors will just sit down with you and sign a contract saying what they will or will not do, they all default to as minimal responsibility as possible. And then they have the guts to jump in on regulation - there’s zero components of the current way these giants do business in AI that is amenable to regulation, at least not in a way that helps you unless you’re an uber-enterprise.
Sentinel Devices | https://www.sentineldevices.com/ | Atlanta, GA | Full-time | Onsite
Sentinel Devices develops "zero-cloud" anomaly detection devices for industrial machinery. Our platform, OTAware, reduces the time users have to spend identifying and tracking down issues in their machines by analyzing machine signals and providing engineers with useful pointers on what's going on, as well as identifying signs of issues that would otherwise go unnoticed. Our unique approach is we do EVERYTHING, including data storage, ML training and decision-making, entirely on embedded devices - we never send data outside of the customer facility. This means that we are an instant buy-in for security-sensitive customers or customers with remote operations, such as defense or oil & gas.
We are ACTIVELY hiring for founding software engineer and AI/ML engineer roles: https://sentinel-devices.breezy.hr/
Also, if you're curious about the roles or the tech feel free to reach out at hello@sentineldevices.com
I sympathize a lot with the headline statement; it boggles my mind on a lot of the data residency/integrity/confidentiality measures taken around massive data silos (as well as the infra teams companies bring to bear to manage, scale and then inevitably publish gospel articles on the web about) when companies could just opt… NOT to collect that data? I really like the model of “It stays on your device, we never see it. At most we get bare-minimum location statistics.” Although I question the assertion that their metrics system won’t be turned against them; seems obvious that anything programmed can be reprogrammed or updated, especially in the modern update-focused age. I don’t think they addressed that beyond a general statement that they took pains to assure that their users won’t ever be spied on. Would be interested in a technical article on that.
Side note, we at Sentinel Devices are taking exactly this “we don’t hold your data” approach for industrial machinery. Think automated AI pipelines that are air-gapped. And we’re hiring! If you’re interested, reach out to hello@sentineldevices.com
Since this is a topic that’s near and dear to my heart, going to quickly plug my own startup that’s doing exclusively this: www.sentineldevices.com. We make industrial equipment smart enough to self-monitor and self-report issues using AI, but we are specialized in making AI that does EVERYTHING (including training) on embedded devices (think something only marginally more powerful than an RPi 4) so we never call out to a server. As the article notes, a big challenge has been making the AI and all support software resilient and able to “bounce back” from random occurrences in industrial environments (which can be extremely unpredictable and unique). Also as the article notes, this pretty handily opens up the Defense market since we don’t need to clear a cloud component.
Feel free to reach out at forrest@sentineldevices.com if you’d like to chat. We’re building and trying to take people on part-time if anyone’s interested in this space! Also if anyone just wants to talk about challenges in the space I’m happy to, HN usually gives really great conversation partners.
Sorry, but this is untrue; action potentials in neurons have extremely complex interplay with each other, including residual “soft” periods and chemically-induced changes in how they fire. Neurons don’t just “fire” or “not fire”, they adaptively change the strength of their firing constantly, unpredictably and continuously.
www.sentineldevices.com
We use machine learning to monitor industrial equipment for signs of faults or failures, and identify in real-time which signals are relevant to the failure/which ones a technician should look at first. The problem we're solving is that when a machine fails unexpectedly, 60% or more of a technician's time is spent just figuring out what was going on and what, specifically, went wrong. We want to cut that time by half or more by having our device be an engineer-in-a-box monitoring the equipment 24/7/365. We're also unique in that we're "zero-cloud" - we do all data collection, storage & processing (yes, even the AI training - not just inference) on-device, on a COTS hardware platform that fits in your hand. The idea is to be truly plug-and-play without having to figure out network infrastructure, and cybersecurity, and data storage costs, etc. etc. Demo video here: https://www.youtube.com/watch?v=FhtLS3UfnPU&feature=youtu.be
We're always interested in pilots; our website is admittedly fairly stealth mode, but if you know someone that works at a factory, they can reach out to forrest.shriver@sentineldevices.com
Hey, I’m a big fan of what you’ve been doing since I stumbled upon it in 2021. Do you have an estimated timeline on when you guys will have that web-app API available, even as a beta? Incorporating a web app with the same ease textualize can be implemented with (or even getting a web app with cross-compatible TUI as a fallback) would be super attractive for the product I’m currently working on.
I guess I’ll take the opportunity to throw in my hat… I’m currently running a very early-stage startup trying to address some of the brain drain and complexity in industrial maintenance. If you’d like to talk/rant about your job and the pain points you have, I’d be all ears! Wouldn’t be a sales call, I’m just trying to talk to people at this point. Feel free to DM me if interested.
I'll take this opportunity to plug my own question to the folks on HN - does anybody know how much protection an LLC generally provides you w.r.t. lawsuits and such? I'm currently looking at doing software consulting for a startup, but indemnification clauses have been touchy for the people I've talked to (and I'm paranoid about losing more than I gain in an unfortunate situation). My lawyer has told me that "piercing the corporate veil" is difficult and an LLC should protect me relatively well, but a good friend told me that I should basically always limit my liability to the amount I've been paid as in his opinion the LLC would be of questionable utility. I know HN isn't the place for legal advice, just wondering if anyone has friendly opinions.
I tried googling this, but the SEO spam is so bad I always converge to the same damn sites. Maybe google search for legal layperson questions is the next big thing?
I agree, but I haven’t found any decent cheat sheet resources discussing problems in all these dimensions. Got any ideas? Open to textbooks as well.
Just curious where you heard Kevin Fu speak about this? I heard about the PowerGuard a while ago but never heard what became of it. If he has talks about the failures they encountered I'd love to hear them, because it's a really cool concept.
I agree with your objective, however it's obvious why Microsoft didn't do this: they wouldn't have been able to make good on their billion-dollar investment in OpenAI/GPT-3, which they REALLY want to justify.
“Those darn Chinese! They’re also launching their projectiles at 45 degree angles! They’re obviously copying the techniques we invented when the laws of physics first came into being!”
That’s you. That’s how you sound.
I have mixed feelings on this article.
On the one hand, I know this is a sentiment that I’ve seen echoed amongst some of my colleagues. Simulation codes just don’t FIT, sometimes, into the neat little box that these mature hardware coprocessors try to put them in. It’s easy enough for ad revenue peeps to adjust their data and models to a new architecture, since the only will they’re obeying is their own and the computational reality is more-or-less whatever they want it to be. With simulations a lot of these assumptions go out the door, because now you can’t just disobey fundamental laws when convenient. If you need to go through a calculation on a single core because you have to evaluate every state sequentially, sometimes that’s all that can be done.
On the other hand, this article seems to be subtly beating the hardware co-design drum, which in my opinion isn’t always a good idea either. One argument could be made that limiting ourselves to one computational approach actually encourages creativity because it encourages some bright young researcher to come up with a new way of looking at things to make the software better fit the hardware. An argument could also be made that co-design is sometimes a bad idea because it might result in things being shoehorned in when they have no place. I’m certainly guilty of experimenting with FPGA implementations which, at the end of the day took so long to compile they would never be useful.
As in all things, I think in this case both “sides” have a little something to offer.
“Emphatically” is a good one.
Well...
1) even if we were all-knowing about what we’ll need in the future, we still don’t know enough about genetics to reliably “push” towards a set goal. Yes we know genetics plays a huge role in things like your body condition, your age, etc., but the degree to which those genes affect - and are affected by - your environment and at what stages they begin working, and where we should alter them, is so far almost completely unknown. Eugenics is grounded in the old idea that much of “you” is just genetics, but we know that idea now to be not just old-fashioned but outright wrong, in that there is a combinatorial number of possibilities enabled by your environment. It’s like calling the codebase of google search the same as the discord codebase because they’re both made of underlying C code (or whatever); technically true but also not true at all.
2)we really don’t have any idea what we should be pushing towards. Maybe right now we think that pushing towards greater strength, for example, would be a good idea, but maybe down the road we discover that one of the genes we made dominant actually makes us very susceptible to some virus strain, overall making the population die at a much higher rate compared to how things would have gone if we had not tried to play god. There are many, many variables you can’t even begin to account for; I find the words “Life finds a way” remarkably suitable to this discussion.
This is a dangerously wrong and even ignorant opinion. There is everything wrong with eugenics, and it’s all contained in the core idea that you have some notion of what is “good” for your children (and correspondingly, further parts of the human race). Other people have already pointed out the “eugenics will be used for political gain” component, which is absolutely true and kills it there. Another issue which I don’t think you appreciate though is the need for genetic diversity and genetic drift. The act of going about systematically pushing everyone’s genes in a certain direction is not only impractical, since we really don’t know enough about genetics to do that reasonably well, it’s also incredibly stupid because we’re getting a temporary “positive” payoff now by potentially screwing up our ability to adapt later. Mistakes are the spice of life, literally; don’t try to get rid of them, you’ll only hurt yourself.
Which lock is it that you’re using? I’m kind of in the market atm.
The water filter question is interesting, would you mind sharing which sites you used for guidance? In this same situation right now and having trouble navigating the waters.
Sorry, I know you wanted to communicate a lot with that paragraph but I gotta ask - what keyboard do you use? I want to avoid carpal tunnel myself but haven’t been able to find any recommendations online.
That’s just direct economic analysis. Graduate students are incredibly valuable to the universities because they generate research which can be used to either directly apply for grants and federal/commercial awards, or which can be used to just add to the university’s reputation (which, in turn, affects their chances of getting a grant or money award, as they have a reputation of expertise). Even if the direct money being paid by these students is zero, I’m sure all universities (not just MIT and Harvard) care deeply about where their graduate students go; that’s a lot of good grey matter going to waste, and universities are in the business of converting grey matter directly into money!
I can elaborate a bit. Most of the large-scale problems are actually as straightforward as “just throw a supercomputer at it”. However, just like when mathematicians say “that’s an implementation detail”, it turns out that actually throwing a supercomputer at the problem is much more difficult to do in practice than merely setting up some shell scripts to run, especially where scientific computing is concerned. For one, there’s usually no concept of “micro services” or “containerized” applications, at least not in my experience. Most of the modern distributed computing practices are actually thrown right out the window when it comes to scientific computing, since the scientists are going to be directly programming distribution schemes via MPI and stuff. The reason is because academic projects don’t have lots of money and need to efficiently use every dollar, and because most of the time distribution schemes really aren’t suitable for the science. You might have one layer where node interactions occur according to some mathematical and physical criteria instead of “load” or some other abstract flag, for example; that’s a bit harder to code for, and it’s much better to have a domain scientist who knows the physics deciding how to decompose the problem, instead of a computer scientist who has no idea of the physics adopting a scheme which ignores the problem entirely. Hence why I said most scientific tools are “bespoke”.
The result is that most of the distributed systems people are moved to a supporting role, where their job is to develop tooling and libraries to allow for better communication between nodes, for example. I’ve also heard of some compsci people being directly integrated into these scientific teams to develop specialized APIs and such in-house, but that’s a bit more rare imo. These are just some examples of how science and compsci intersect; for example here’s one group I know of: https://www.ornl.gov/group/dcs
I’m biased as my work is primarily focused on large-scale distributed physics simulations, and incorporating machine learning into these. As a result, I treat ML very much as a means to an end.
Of course with the caveat that your situation is unique to you so I can’t give any definitive answers, I would think long and hard before jumping on the ML hype train. In my experience, it doesn’t pay to follow the trend; you’ve either gotta be first or you gotta be unique. Now that’s not to say that doing ML work work will only be restricted to a select few which you aren’t a part of, but myself and a few others are wary that the ML hype train (at least as far as deep learning is concerned) might be passing. The days of the AI labs paying million-dollar bonuses are nearly gone, unless (and someone can correct me if I’m wrong) you’ve got an alternate skill set they’re looking for. Of course, that doesn’t mean there aren’t plenty of people and businesses who would need CRUD-type ML setups; with your experience in databases I imagine that could be a unique angle to attack it from. Whether it’s a good idea to try and pivot into a career using ML really depends on your specific situation and the opportunities therein; to get more solid advice I would ask a trusted colleague or mentor, and would not consult people online, even if they are from HN.
For my PERSONAL opinion: I can’t speak to what is normally done in other parts of distributed systems, since scientific tools are usually bespoke and don’t use the same set of approaches as commercial products. However, just thinking about it from an outsiders POV, it seems to me like focusing more on distributes systems would be a winning combination. I don’t think computers will advance enough in the next 30 years that the need for distributed data and compute management skills will go away; hell with IoT you might be looking at a boom down that career path. From my perspective it’s only upside if you focus on expanding your skilllset in these areas; if ML continues to thrive there’ll most definitely be a need for distributed systems to run these models on. And if an AI winter hits, you’ll have a solid set of core skills to fall back on which I don’t imagine will go out if favor anytime soon. Those are just my two cents though, of course YMMV.
I’m unsure how to feel about this. On the one hand, I can most definitely sympathize with feelings that a name, no matter how inoffensive to some groups, might be deeply offensive to others. However, in my mind there’s the question of standardization to be considered. Science (unlike the software world) can’t change every few years (at least it shouldn’t). Some of these definitions have become so ingrained in existing literature and cultural knowledge over the decades that you have to ask what happens if you suddenly decide to change something; you have to wonder what effects it can have on a field if something is suddenly and forcefully renamed (students can get confused, older but sill quite valuable reference texts can be made confusing to read).
I also had issue with how the article decided to present its opening argument, by saying that an African American felt disturbed when they went to “noose” lizards in a group of white people. While I absolutely agree that this would be in bad taste if they were at all implying doing this same thing to people, the context I got is that it was a purely academic term derived from the shape of the tool they’re using, and NOT at all oriented towards the act itself or the specific things they’re trying to catch. To me that sounds very similar to someone saying that programmers referring to “killing the child” is offensive, when everyone knows the act is not related at all to the death of actual children and is not trying to make light of situations where that happens. It’s a historical and technically functional reference, nothing more.
To be clear, I don’t feel that all of these causes are wrong. I’m just worried that in our haste to make symbolic changes, we’re pushing towards alterations that only have an effect years later when everyone that would be able to explain the changes is dead. I’m also worried that these very symbolic changes will be used as a band-aid and aren’t addressing the deeper issues at play here, and that they might even distract from what actually needs to be done.
That literally makes no sense. How exactly would a quantum computer assist us in simulating problems like drug transport and binding? (quite a few problems are purely empirically solved right now) How exactly is biology massively parallel? I suppose you could argue a cell is just a weird form of SIMD but that is a vast oversimplification. We’ve also been doing parallelism pretty much since computers were invented (and even before that, assembly lines were a thing) it’s just that now we really care cause our one trick of shoving transistors in doesn’t work so hot anymore.