I wonder how long it would take a "normal" coding prompt to go thru a "1T-2T" models on a medium performance consumer desktop.
Hours? Days?
HN user
Stop blocking login from user agents which are not using whatwg cartel web engines, thx.
Is your "security" provider making love with gogol and its whatwg cartel friends or do I have active parasites on my internet lines?
Well, HN comments are kind of dead: too many AI bots and a abused "karma" system. Not allowed to disagree or provide alternatives.
Sorta no point in argumenting.
Still, we can use news and comments as a channel of communication and publishing.
And support email addresses with IP literals: mailbox@[x.x.x.x] mailbox@[ipv6:...] for those who are self-hosted and not paying the DNS mob, like myself. You can filter hard incoming emails with IP matching from the TCP connection to the content of the 'from' related headers, this is way stronger than SPF, thx.
I wonder how long it would take a "normal" coding prompt to go thru a "1T-2T" models on a medium performance consumer desktop.
Hours? Days?
I wish that, but for high performance RISC-V micro-achitectures.
Once I have something satisfying, I'll post the "RFC-like" draft and preliminary reference toolchain on HN (well, if I can since HN is going 'whatwg cartel' web engines only then I may be blocked for good). For the moment, I am 'testing' all that while writting my own wayland compositor for linux (that's why I will try to build mesa AMD vulkan driver for this format... or stick to wl_shm in the end).
Well, for a new executable and dynamic library format, the computer language syntax complexity and its runtime complexity do matter A LOT, c++ is massive pain, and since, if I am not mistaken, sqlite is going microsoft rust, that's bad omens since microsoft rust syntax seems to be now as brain damaged than c++ syntax... but with an even worse runtime (unless all that is not actually true, I have not checked). That said, ISO f*cked up C with things like '__thread' which requires some level of OS support.
With the combination of the permanent torrents of software and hardware flaws, instrinsics to their complexity (and bad strategic choices), there is only one reasonable, honest, brutal realization:
information system security is a fantasy, it does not exist.
You can only mitigate in a best effort fashion.
Don't have a choice, will probably have to go open-weights models as currently, "AI" is gated using 'whatwg cartel' web engines.
In the light of this, I am mechanically a proponent of very good open weights models, which I can download (for instance on on bittorrent) and run, slowly (the price), on local hardware.
That would be for coding.
If china puts its AI models on the same ground than US capital investment funds and big tech financial support (aka Big Tech international finance), they will very probably lose everything (know how, ML and inference infrastructures).
linux ELF code is loading the ELF loader only. $ORIGIN may not be a good idea since that would add more ELF complexity to the kernel.
If we are honest with ourself, ELF is the core of the issue: for executables and dynamic libraries we _now_ know it is severely obsolete on modern hardware architectures.
I am currently using my own format, excrutiatingly simple, no loader, basically a program segment, "userland syscalls" with hardware CPU synchronization. I do wrap the executables into ELF capsules to run them transparently. So simple a small RFC will be enough.
With that, I discovered that the hard part is c++ and other similar languages which are very expensive in runtime infrastructure and linking complexity. I would need to build a mesa vulkan driver with that format, and it seems the blocker is c++ (and similar language namely with grotesque and absurd syntax complexity). Thx to valve to have removed a lot of c++, for less c++... would have been much better if plain and simple C.
Sanity and common sense are starting to fight back?
You are an odd one: we all know that c++ syntax complexity is beyond salvation : it has reach such an absurd and grotesque level, we are in mental pathology realm: Rube Goldberg Machine accute syndrome. And it is not even a matter of argumentation, unless being of accutely bad faith.
It seems you are failing to see that this is this very syntax complexity which makes most of this toxicity.
And having a modern c++ transplier to plain and simple C, would re-open the door to alternative and 'real-life' small/medium compilers, not like those vendor/developer-locked very few options we have today.
Somebody did it for microsoft rust. Would just be a good idea for nowadays c++.
EU have to go on the technical ground and force interop with small, but able to do a good enough job, stable in time, protocols and file formats.
A good compromise: subsets of protocols and file formats. For instance, for the web, noscript/basic HTML interop to break their whatwg cartel (nearly all web services were provided with basic HTML forms only a few years back). Another example, a curated subset of PDF (of course excluding the brain damaged abomination which is 'forms with javascript in PDF'). That way you retain interop with absolutely everbody while keeping the door open for real-life and small alternatives. A warning though: expect Big Tech to try everything to break it, don't forget they are serial offenders and you don't know how deep the rabbit hole goes (shadow-paid hackers? very $$$ oriented lobbying?).
This means you can access and interop with real-life alternative small web engines (past, present and future). It means you can render in a good enough fashion a PDF file (the only improvement PDF does require is full replacement for a near iso-functional file format, but fully cleaned up and rationalized based on PDF experience and simplified from a parsing point of view).
A note for the web, some trash human beings will try to scare EU with 'security' and lure them into the software they control (usually the whatwg cartel). 'security' of an online service is where 99.9% of the work happens: it requires a permanent team of people who can be trusted (hard) close to network operators/data centers/authorities which monitors and hunts. That cannot be 'small' in any capacity, and the following must be presumed: they will fail more than once and more often than you think: this is only a best effort because 'information system security' is a fantasy and does not exist.
Ofc source, super simple file formats must be supported: for instance the ubiquitous utf8 text file or the PNG image (with probably only a subset of pixel formats).
And for AI, we could have a 'curl' web API with public tokens (probably severely rate limited), or with a private token registration process, but which can be done with 'physical mail' (yep, you read it well)', or noscript/basic HTML browsers with optionally combined with a self-hosted email server not paying for DNS (email address with IP literals, which gmail.com is careful at not supporting it even though it is stronger than SPF.....), aka small tech. Since we have IPv6 almost everywhere in my EU country, you could get a private token tied to one and only one IPv6 address (aka the real symmetric internet).
A good test is to try to port significant from c++ to plain and simple C using "AI". Maybe it can retain some of the semantics for the original code.
There is a more brutal option: AI could help write a c++ to plain and simple C transpiler... then we would lose the semantics of the original program, but free ourself from the very few toxic and real-life c++ compilers out there.
FOSS is far from enough anymore.
_LEAN_ FOSS, including the SDK then the computer languages too.
All computer languages with an ultra-complex syntax are excluded de facto.
Then there is the stability in time.
developer/vendor lock-in on software, planned obsolescence, are much more common in FOSS nowdays.
Do you know of coding inference models I can access with curl? (namely with public access tokens, probably severely rate limited).
Or some coding inference models I can access with a noscript/basic HTML browser? (namely basic HTML forms)
If those inference models are still gated by whatwg cartel web engines, I will have to run full blown coding frontier models locally (if I want to have a chance at getting quality code). It is going to be very slow, and even slower while I am developping "prompt templates".
I did say "quality code", because I did ask some people already to generate classic and basic code paths using AIs, all were quite disappointing. That said, it was millions of years ago (in AI improvement time), namely a few months ago in human time :)
Zero usage of LLM (assembly coding, and sometimes plain and simple C).
I cannot access them to test if they would be of any help in my coding use cases.
You tell me once we get some inference access with a web API with public tokens (probably severely rate limited), or with full interop on noscript/basic HTML.
I may have to run locally open weight coding frontier models (slooooooow).
Yep, I guess we are many to see what's happening with RISC-V.
The most risk of cruft accumutation is in RVA... which is pursuing some level x86-64/ARM hardware compatibility.
That said intel APX/AVX10.2 is RISC-V for x86-64...
Everything pushing forward RISC-V is a good thing (this time I get it right...)
I code RISC-V assembly almost everyday, beyond the major point that it is a NON-IP-LOCKED ISA (unlike arm and x86-64), it feels like it does 'sweet spot' nearly all the time. Namely, I am more into binary specifications which means, if RISC-V is zapped one day, we still have some RISC-V byte code and port to an IP-LOCKED ISA is reasonable.
The hard part: _really performant_ micro-architectures for server/desktop/embedded/mobile on latest silicon process.
The harder part: getting much binary-only 'critical' software running there (for instance desktop video games).
And the super hard part: big mistakes _will be made_, and it is going to hurt ooofely.
It is not, the 'static' part is just some sugar to remove a bit more information from the internet traffic generated by such communication protocol. The software development part is mostly to mitigate what's well know with current planned obsolescence abuse from many software actors (and move some risks from corpos to people).
I2P is obviously much less a fantasy than 'information system security' but is very creative and dependent on the current internet trends: you can encode in clear text in video games chat/socials, have fake streaming services (asymetric), fake video games, etc, etc. I heard 'rumors' a long time ago about about 'special commands' encoded in the silicon of some carrier grade routers, were they true?
I2P is a rabbit hole which can go, REALLY, REALLY DEEP. Basically, if the guys are carefull, the 'services' or experienced hacking team would need to know who to look at.
But here, nope, just something to help 'privacy', but as I said, with some limited impact on mass-monitoring and "John Doh" ("hacker only on Sundays").
The "less worse scenario of internet communication privacy" is direct IPv6(coze IPv4 with custom port redirections is messy, local custom IPv6 built on ISP provided prefix is much less messy) with custom UDP(or TCP, since its usage is broader than UDP then looks more "normal") ports, E2E encryption with non-public pre-shared keys. If you are competent enough, you would have custom modifications on some standard crypto (usually session pre-shared masks, which generation code is pre-shared) with a super simplified video/audio/text protocol (probably a super lean SIP or XMPP). Ofc, no clear protocol negotiations of any kind, the traffic is immediatly encrypted. If you really want to push it, your protocol should be isochronous and have fixed size data frames ("static", with probably an exception upon tranfering really big data files, unless you are really patient). Of course, everything coded in assembly on various ISAs, optional simple code generators coded in assembly themselves without abuse of any assembler pre-processor, which means the usage of some common binary specifications).
This should give you a rather "good" protection against mass-monitoring, and against John Doh, "hacker only on Sundays".
Yeah, and we did not talk about your "corpo" hardware, your system software built with silent backdoor inject... ahem... I meant "compilers".
And in the end with all that, if the 'services' or some experienced hacking team really want to watch _you_, don't worry, they will: they don't need your password or any of your crypto keys/data.
But on the overall, based of the documents and information provided here, and what I think I managed to understand about them: it is not surprising some want to start the deprecation of C from RVA (some would even start its deprecation from the core specs to free some machine instruction encoding space).
It is a bit like compilers: 90% of the complexity/size is for 30% of the code optimization speed gain, which has mostly a significant and pertinent impact in specific work loads on our modern hardware. Basically it is "better" (since you get rid of 90% of the complexity/size of compilers) to combine pure assembly coding with less complex/bigger compilers ("10%") (and look at dav[12]d, not to mention ffmpeg, here the main issue is feature creep of ISO C new things and GCC extensions... and abuse of the nasm pre-preprocessor). For instance, with bzip2, with cproc/qbe (zero assembly though), I get 70% of the speed of modern gcc/clang. But what's really worrying: I talked (I think it was here on HN a long time ago) to one of TheHeavyThing assembly coders, and he said that with a brutal and naive hand-compilation of the corresponding C code, he got 15-20% faster than the best compiler could provide at the time on gzip compression. Something is up here, and I already did mention dav[1]d and ffmpeg.
We are in a 101 case of 'compromise', and here, it seems the more I read about it the more I think we are facing a case of 'benefits are not worth it in the end', well the apparent 'competent' people who are in the sauce seems to agree more and more about it, through the infos I have been getting from HN.
As for the 'compilation' benchmarks provided here, we know gcc/clang are heavily x86-64 oriented, namely goes into optimizations which should favor 'C' like machine instructions (destination register same than some source register), so the 1-2% "real" speed increase from some documents provided here sounds more like the real neutral/overall thing to me.
This is another way to say they are scared ofe competition and want to scrap them.
I thought so.
But I want that :)
Huh?
There are RISC-V "performant" implementations on the best TSMC silicon process like x86-64 and aarch64?
Many here wish for the non IP-locked RISC-V to get performant micro-achitectures for embedded/server/desktop/mobile that on the latest silicon process (without that, you can have a very good micro-architecture, that won't probably make the difference).
If some RISC-V high performance CPU manufacturers are being bought by big hardware actors: either they are scared of its competition and want to scrap it or they want to be part of RISC-V.
Well, I said that because based on the documents provided here, it seems there are key people considering its removal even from the specs.
From my point of view, just do like all the others: clearly deprecate it, namely say that "from RVAx, don't create new machine code with the C extension". It is like in the linux kernel, it will then be removed very far in the future. But RISC-V is all about the far future, it can only be better to fix it asap.
After reading the comments and documents provided here, I am the first to be suprised by how much doing 'performant C' is not that easy and has a significant hardware cost. "arm removing thumb" should have been a strong signal.
Again, I am coding rv64 assembly almost every day, a good part could be C-ized to shrink text size, but based on various numbers provided here, why bother, better keep the 'R' of RISC as faithfull to its goal than anything else.
Thx for the github.com link, I can have a look at it.
mitigations do slow all CPUs, even those without the vulns.
You need alternative linux kernels with their modules.
Maybe the best approach is to remove C from RVA while keeping it around in the specs for niche applications where text size _really_ matters (with current silicon processes, I wonder how weird those niche applications have to be to require C). But it seems some would remove C even from the specs to free some ISA space. If arm removed thumb...
Don't worry, I would know if the assembler is producing C machine instructions, my rv64 interpreter on x86_64 does not support the C instructions at all.
Cannot access the web site: whatwg cartel web engine only.
Is antjs coded in plain and simple C?
You have to recompile the kernel to remove some mitigations about "indirect branching", namely you must have alternative windows kernels.
Do you have those windows kernels?
English is not my native language and I wrote a bit too fast the message: I wanted to say that "everything pushing forward RISC-V is good".
I code RISC-V assembly, I don't use C machine instructions (I don't even use the pseudo-instructions, ABI register names and dodge nearly all ISA extensions, I try to stick to core as much as I can). I run my code on x86_64 linux with a small interpreter written in x86_64 assembly (thx to the 'R' in RISC).
I wonder if there are some 'broad and not niche, real-life' speed benchmark numbers to show how much C machine instructions are worth.
For the moment, I see those C machine instructions more as a marketing extension to match their arm equivalent: you know, for those key deciding people who care more about the amount of features and not their contextual pertinent usage.
To say an ISA is "good" is related to some set of technical sweet spots based on compromises based on projected usages.
RVA from my point of view is mostly preparing RISC-V hardware for some level of x86_64/arm compatibility.
I wonder if there are RISC-V implementations using the latest silicon process from TSMC.
"Free Software" is not enough anymore.
We need _LEAN_ free software, including the SDK, then also the syntaxes of the used computer languages.