HN user

andthenwhat

22 karma
Posts0
Comments7
View on HN
No posts found.

not OP but it does seem like you're nitpicking details instead of engaging with what seems to be the intent of the response: AR/VR has come an incredible distance since DK1, the last 10 years have seen it go from a barely-discussed completely unavailable/fringe dev-kit-only technology to being an clearly viable spectrum of mass-market products.

edited: grammar. still feel like I've failed to produce readable english, but I'm giving up

For the image generation (or even indexing with the CLIP interrogator) side of things, recommend just installing the AUTOMATIC1111 github repo (https://github.com/AUTOMATIC1111/stable-diffusion-webui), it's a web ui with pretty much every variant of stable diffusion you could want to try out, like txt2img, img2img, inpainting (both textual inversion and dream booth), outpainting, style customization, clip interrogation, etc. Most importantly, there are about 1000 youtube tutorials on how to do each of these things with it, so you can pick your interest areas and just try it out without having to understand all the details first.

From there, if you're interested in how it works, I highly recommend the last 4 videos on Jeremy Howard's youtube channel: https://www.youtube.com/user/howardjeremyp/videos

He's currently teaching a class on stable diffusion from the ground up and these lectures give a really good introduction to how it all works.

That goal might already have been somewhat lost years ago. The video almost everyone after the editors watch has been re-encoded at least once (more likely twice) with one or more different codecs that definitely change the character of the video from that of the source. Any same source content will already look noticeably different depending one which service or provider you’re using, since they each have their own opinionated video content encoding and distribution pipelines. If you’re talking about archival storage in the sense of preserving masters, then I definitely agree and the film grain removal should be disabled in the encoder, so the decoder-side synthesis won’t happen. Thankfully, that’s very easy to do (for example, libaom-av1 encoder in ffmpeg supports denoise-noise-level parameter set to zero to disable those scary parts.)

How do you avoid the problem of using data-trained algorithms to provide measurements (in that they can’t)?

The idea of classifying or localizing detections with deep learning (or any data-trained approach) seems totally reasonable, since it’s clearly making an inference, and that should be clear to the human user. Enhancing or gap-filling the measurements with data-trained approaches would turn the measurements themselves into inference, which seems in opposition of the diagnostic goals (algorithm in-painting non-sensed information learned from the training set)

Agree. I thought I was clicking on an article about screen-space reflections in 3D rendering, although CGI and Perl were around long before that so they win the race for that acronym.