HN user

eightails

317 karma
Posts1
Comments98
View on HN

For sure, and even training is very doable on consumer hardware these days. Techniques like Dreambooth and LoRA have dramatically lowered the compute cost of finetuning large models on specific concepts. A recent GPU can train Stable Diffusion models on a concept using LoRA in < 30 minutes.

I guess that makes sense. I use Mullvad as well, and anecdotally it seemed to have a similar rate of blacklisted endpoints to Nord. I guess in Nord's case maybe its ubiquity is the problem.

Deja vu, I have been in this place before.

It's amazing how circular the responses in this comment thread are getting. There appears to be disagreement on the surface, but in reality almost everyone is presenting a view which is at least compatible with each other's -- if not in direct agreement.

I think it's been raised in the context of a fair few sightings where there's supposedly been a very large craft moving silently and relatively slowly, e.g. the 2000 Illinois sightings (referenced in a great song by Sufjan Stevens). I was looking into this one a few years ago and found references to a private company testing blimp platforms for military purposes around that period, although I can't find it off the top of my head.

More recently, the object spotted hovering off Hawaii last year that resulted in fighters scrambling was also proposed to be a modern balloon-based drone, of which there are a few currently being developed.

Edit: here's a source arguing that the Illinois sighting was a regular advertising blimp

https://skeptoid.com/episodes/4435

And still lower than other laser systems using inertial confinement. If I'm reading NIF's recent paper right, they're claiming Q_a of above ~1.4, or Q = ~0.28

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8791836/

Edit: Wikipedia suggests that they were using a different method to calculate Q, only measuring the power input to plasma vs output from fusion, not including system losses. So that figure is probably not directly comparable.

the NIF used ~477 MJ of electrical energy to get ~1.8 MJ of energy into the target to create ~1.3 MJ of fusion energy

https://en.wikipedia.org/wiki/National_Ignition_Facility#Bur...

But, of course, I'm overlooking something here. Because if you take the same portrait at 50mm and with, say, 20mm, it's not just the focal length of the camera that differs. What also differs is the position of each camera. The 50mm camera will be positioned further away from the subject, whereas the 20mm camera has to be positioned much closer to achieve the same "shot".

Yep, totally.

Perhaps it helps that the vehicle moves? That is, after all, very close to having the same scene photographed by cameras positioned at different distances.

I think you're right, they must be taking advantage of this to get the kind of results they are getting. That point cloud footage is impressive, it's hard to imagine getting that kind of detail and accuracy just from individual 2d stills.

Maybe this also gives some insight into the situations where the system seems to struggle. When moving forward in a straight line, objects in the peripheral will shift noticeably in relative size, position and orientation within the frame, whereas objects directly in front will only change in size, not position or orientation. You can see this effect just by moving your head back and forth.

So it might be that the net has less information to go on when considering objects stationary directly in or slightly adjacent to the vehicles path -- which seems to be one of the scenarios where it makes mistakes in the real world, e.g. with stationary emergency vehicles. I'm just speculating here though.

Also, among the front-facing cameras, the two outermost are at least a few centimeters apart. I haven't measured it, but it looks like a distance not unlike between a human's eyes [0]. Maybe that's already enough?

Maybe. The distance between the cameras is pretty small from memory, less than in human eyes I would say. It would also only work over a smaller section of the forward view due to the difference in focal length between the cams. I can't help but think that if they really wanted to take advantage of binocular vision, they would have used more optimal hardware. So I guess that implies that the engineers are confident that what they have should be sufficient, one way or another.

I'm just wondering if using cameras that are close to each other, but use different focal lengths, doesn't give the same results

I can see why it might seem that way intuitively, but different focal lengths won't give any additional information about depth, just the potential for more detail. If no other parameters change, an increase in focal length is effectively the same as just cropping in from a wider FOV. Other things like depth of field will only change if e.g. the distance between the subject and camera are changed as well.

The additional depth information provided by binocular vision comes from parallax [0].

Also, wouldn't turning a multitude of views into a 3D map require a neural net anyway?

Not necessarily, you can just use geometry [1]. Stereo vision algorithms have been around since the 80s or earlier [2]. That said, machine learning also works and is probably much faster. Either way the results should in theory be superior to monocular depth perception through ML, since additional information is being provided.

It seems to me that this is how modern phones are doing background removal: The lenses are very close to each other, very unlike the human eye. But they have different focal lengths, so depth can be estimated based on the diff between the images caused by the different focal lengths.

Like I said, there isn't any difference when changing focal length other than 'zooming'. There's no further depth information to get, except for a tiny parallax difference I suppose.

Emulation of background blur can certainly be done with just one camera through ML, and I assume this is the standard way of doing things although implementations probably vary. Some phones also use time-of-flight sensors, and Google uses a specialised kind of AF photosite to assist their single sensor -- again, taking advantage of parallax [3]. Unfortunately I don't think the Tesla sensors have any such PDAF pixels.

This is also why portrait modes often get small things wrong, and don't blur certain objects (e.g. hair) properly. Obviously such mistakes are acceptable in a phone camera, less so in an autonomous car.

And those illusions work even though humans actually have an advantage over cheap fixed-focus cameras, in that focusing the lens on the object itself gives an indication of the object's distance

If you're referring to differences in depth of field when comparing a near vs far focus plane, yeah that information certainly can be used to aid depth perception. Panasonic does this with their DFD (depth-from-defocus) system [4]. As you say though, not practical for Tesla cameras.

[0] https://en.wikipedia.org/wiki/Binocular_disparity [1] https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.36... [2] https://www.ri.cmu.edu/pub_files/pub3/lucas_bruce_d_1981_2/l... [3] https://ai.googleblog.com/2017/10/portrait-mode-on-pixel-2-a... [4] https://www.dpreview.com/articles/0171197083/coming-into-foc...

Sure, but they're not getting that 3d map from binocular vision. The forward camera sensors are within a few mm of each other and different focal lengths.

And the tweet thread you linked confirms it's a ML depth map:

Well, the cars actually have a depth perceiving net inside indeed.

My speculation was that a binocular system might be less prone to error than the current net.

One question I've always had about Tesla's sensor approach: why not use binocular forward facing vision? Seems like it would be a simple and cheap way to get reliable depth maps, which might help performance in the situations which currently challenge the ML. Detecting whether a stationary object (emergency vehicle or child or whatever) is part of the background would be a lot easier with an accurate depth map, or so it seems to me.

Plus using the same cameras would help prevent the issues with sensor fusion of the radar described by Tesla due to the low resolution of the radar.

I know the b-pillar cameras exist, but I don't think their FOV covers the entire forward view, and I don't think they have the same resolution as the main forward cameras (partly due to wide FOV).

I'd love to hear why I'm wrong though.

Sounds similar to experimental designs by Rolls Royce for a diesel rotary during the 60s, in that it uses two separate parallel shafts and rotors for intake/compression and ignition/exhaust.

In the case of the Rolls-Royce Wankel Diesel, the fuel-air mixture is first compressed by the lower rotary, and the output of that engine (which would be like the exhaust valve of a conventional rotary) sends the compressed diesel/air mixture to the intake of the smaller upper rotary engine, where it’s compressed to ignite like a regular diesel engine.

Seems like it had the same issues that have always plagued rotaries, primarily with apex seals.

https://jalopnik.com/this-might-be-the-weirdest-engine-rolls...

https://youtu.be/1pDjwaqU0dU

It seems clear to me that there is an optimal starting word, but that the best second word has to depend on the info you gain from the first.

Definitely. The likelihood of a letter appearing in a given place changes depending on the letters around it. A Q will almost always be followed with a U, for example.

I wrote a script yesterday which spits out the relative probabilities of possible letters in each unknown position, given the current known/excluded letters -- it was interesting to see the effect in action.

Totally agree.

I feel like this speaks to a general reluctance to accept uncertainty and shades of gray surrounding a topic. It often seems to me that many people prefer to have things neatly sorted into binary categories, supporting them either completely or not at all, when most topics resist this kind of neat division if given more than a cursory glance.

Articles refusing to include any nuanced discussion of a point and simply stating an arbitrary position as proven fact just exacerbates the problem.

It's not too hard to make a connection with the often-polarising and vitriolic nature of discussions online and the rise of misinformation.

Yeah that's likely true. I had chosen MP3 for reasons of total compatability and 'good enough' compression, but it seems like almost all music players have good support for those formats as well. Maybe at some point I'll switch over.

I agree. I was using almost entirely lossless until five years ago or so, when I did some ABX tests (there's a good foobar plugin for those interested) and realised that 256+ kbps LAME was indistinguishable from lossless 24/96 to my ears -- even when actively comparing the two in a quiet room, which obviously is very different to just listening casually while commuting etc.

I kept the lossless files for archival purposes but everything on my phone/laptop is ~320 LAME VBR.

That said, soon after that I switched to Spotify and only rarely listen to my own files now. The convenience and ability to discover new music just doesn't compare.

Not appreciably different, that was the whole logic behind my thinking. In my case I have electric heaters, so doesn't make a difference compared to a gpu -- they're both effectively 100% thermally efficient.

I did this last winter. Mining 24/7 on my personal rig noticeably took a bit of load off the heater and made a modest amount of ethereum. I stopped once the weather warmed as it seemed like a waste of energy.

I suppose an optimist would say that success, even at a tiny scale, would be a step in the right direction. But you're right, it could well be possible that the effect doesn't scale up.

I think gram-scale probes are definitely being considered -- though not with Alcubierre drives obviously lol. Breakthrough Starshot think they can transmit back from Alpha Centauri at 2.6-15 baud per watt by using their light sail as a laser reflector [0]. Pretty crazy.

[0] https://en.wikipedia.org/wiki/Breakthrough_Starshot#Laser_da...

The issue seems to be that the alcubierre drive requires things like external negative energy

Yes that's traditionally been the stumbling block. But the whole point of the linked article is that they have predicted a way to meet this requirement:

a micro/nano-scale structure has been discovered that predicts negative energy density distribution that closely matches requirements for the Alcubierre metric

Obviously it's still a very very long way from a practical application re. Alcubierre (if such a thing is possible), but it's certainly an intriguing result if correct.

From a quick flick through the paper, the above commenter seems to be correct in saying that they haven't yet completed a practical experiment to confirm. So nothing more than a simulated result at this stage.

I have a very weird requirement when it comes to reading digital content. It has to be 100% distraction free because otherwise I end up getting sucked into Netflix or HN

Not exactly a radical recommendation or anything, but have you tried a dedicated e-reader? No distractions possible that way, and you get the all the usual benefits -- e-ink screen, weeks-long battery, etc. I enjoy mine.

With digital books I often finish a book and still don't know the author because the only time I've seen the cover was when I was at page one

Kobos (maybe Kindles as well?) can display the cover of the current novel while they're sleeping.

In a way it reminds me of being engrossed in reading, in that there's an almost complete suspension of awareness of the outside world in favour of the fictional one. The fact that RS users report writing detailed scripts seems to support this:

Before I plan on shifting, I write myself a script in the notes app on my phone, in which I plan exactly what happens in the desired reality. This makes it easier to visualize exactly what I want to happen

I suppose one difference is that in reading, you (or I, at least) tend to also dissociate the self somewhat in favour of the main character in the novel, whereas RS users seem to retain their sense of self.

Following this train of thought, I wonder if reading could itself be considered a kind of meditation?

This sort of thing is so interesting from a psychological perspective, is a shame it has to come laden with bogus rationalisations via the occult or pop-quantum physics, which just provides an avenue for easy dismissal of the effect.

As an open-source alternative, Joplin could be worth a look. It's a little different in terms of features, missing a lot of the linking which is kind of key to Obsidian's appeal, so maybe it's closer to Evernote. Still, they're fundamentally the same in terms of being a note-taking app built on folders of markdown files. It's quite actively maintained and improved too.