That dependency is now gone, thank you for the review
HN user
flatroze
Everything has its limits
Could you please open issue on github providing the output that you get in the terminal?
It sounds like a great add-on, I have to check it out to see what it does to remote assets and how it works with asynchronously loaded assets.
I'd say since monolith produces a plaintext document it lets you edit things easier if needed.
JS can be removed from the final document using the -j flag. HTML Files can also be grepped for content, unlike PDFs.
The compile time is rather long as well, I'm looking into ways of reducing the amount of dependencies.
Thank you for the kind words.
It will evolve into a reliable tool in a couple weeks and it should eventually work for embedding everything, including things like web fonts and @url()'s within CSS. If anything doesn't work, please open an issue, I have plenty of time to work on it.
Which license would you recommend to release this software under to reach the broad adoption yet permissive terms, if not the Unlicense?
It is done, option -i in the latest version (2.0.3) now replaces all src="..." attributes with src="<data URL for a transparent PNG pixel>" within IMG tags.
CSS imports are covered by converting .css files into data URLs, later I will parse those and embed resources found within stylesheets as well.
It's a valid question. It seems to me users tend to trust things which have certain level of popularity and reputation associated with them.
I personally prefer to hope for the worst. This way when nothing happens I feel extra lucky, and if bad things do happen, I feel proud of being ready for it.
Thank you for the heads up, I'll test it and enhance to preserve styles better.
Thanks, I'll look into making it work with pipes or some other way to interact with headless browsers.
Thank you for reminding me, I need to set action="" to be an absolute path when the page is saved.
upd: Done, now forms get their action="/submit" converted into action="https://website.com/submit" when the page is saved.
Well, at least it's not called iSuck.
Thanks! Pictures should work, I'll check more tags first thing tomorrow when I start working on improving it.
I use youtube-dl for youtube and other popular web services myself. Embedding a video source as a data URL could in theory work, but it'd be quite a long base64 line. Also, editing .html files with tens or hundreds of megabytes of base64 in them would perhaps be less than convenient.
That's it in the nutshell!
It seems to work for basic pages quite well, I think that lazy load will work for most pages as long as the JavaScript is embedded (no -j flag provided) and the Internet connection is on. It saves what's there when the page is loaded, the rest is a gamble since every website implements infinite scroll differently.
Authentication is another tricky part -- it's different for every browser. I will try to convert it into a web extension of sorts, so that pages could be saved directly from the browser while the user is authenticated.
Uh, I'm sorry. Please feel free to open PRs to that repo if you have anything to add or improve.
It for sure would help with those SPA websites that get their DOM fully generated by JS. A web extension that saves the current DOM tree as HTML would perhaps do a better job, especially when it comes to resources which require some web-based authentication.
Ah, I remember using something like that. I thought that tool was saving it into one .html file, but data URLs didn't exist back then, so creating directories alongside with HTML files was the only option to "replicate" a web resource, now I understand exactly what you were talking about. I'll do some more digging around and implement that in the nearest future. I may need to make all the requests async first to make sure that saving one resource with decent depth won't take too long.
This could also be an interesting alternative to PDF, especially with web fonts embedded as data URLs.
We'll bring it back, don't you worry!
The idea is almost identical, yet saving as .webarchive is only supported by Safari, and it's also not a plaintext format, hence can't be edited as easily.
Oh, thank you kindly.
That's an interesting question. I think it depends on how the given modal is implemented, but closing them should technically work (unless the page is saved with JavaScript code removed [-j flag]). Those notifications can easily be removed from the saved file using any text editor, should be pretty easy if you know how to edit HTML code. I don't think removing it would violate anything since "this website" will no longer really be a website but rather a local document at that point.
I think there's an issue with opening a tar file, e.g. if sent to someone who needs to view the document but isn't techy.
It seems to me that having one file that any browser can easily open (and not require Internet connection to view) is a big advantage over having a directory with assets alongside the .html file. It may be one of those things that make things easier yet nobody really complains about how things are usually done when the page gets saved. I hope more browsers add support for saving pages as MHTML in the nearest future so that we wouldn't need tools like this one.
Good points, thank you for the review. I'll work on enhancing the readme file to be more informative.
Apparently everybody knew about MHTML but me Ü I'm going to look into that format and see if I could enhance monolith to output proper MHTML, among other additions and improvements. Thank you for the info!
Thank you! It's pretty straight-forward: this program just retrieves assets and converts them into data-URLs (data:...), then replaces the original href/src attribute value, so in case with the same image being linked multiple times, monolith will for sure bloat the output with the same base64 data, correct. I haven't looked into MHTMTL, ashamed to admit it's the first time I'm hearing about that format. I need to do some research, maybe I could improve monolith to overcome issues related to file size, thank you for the tip!
And about Rust: I think you're way ahead of me here as well, this is my first Rust program. If you're talking about it embedding some debug info into the binary which may include things like /home/dataflow then perhaps there's a compiler option for cargo or a way to strip the binary after it's compiled. ¯\_(ツ)_/¯ Sorry, that's the best I can tell at the moment.
Thank you! I'll add it as an issue, since it could definitely be useful for "archiving" certain resources more than 1 level deep. Do you remember the name of that tool by any chance?
"And so it happens that the person who reads a great deal — that is to say, almost the whole day, and recreates himself by spending the intervals in thoughtless diversion, gradually loses the ability to think for himself; just as a man who is always riding at last forgets how to walk."
Source: https://ebooks.adelaide.edu.au/s/schopenhauer/arthur/essays/...