HN user

TikiTDO

84 karma
Posts0
Comments44
View on HN
No posts found.

Looking at that paper, they appear to be saying that 6.7B is where the problem becomes so intense that no single quantization method can keep up. From what I gather, the paper claims that such outliers start occur down to 125M param models, then at around 1.3B they begin to affect the FFN, and at around 6.7B is when the issue really starts to become apparent because "100% of layers use the same dimension for outliers."

So while you obviously wouldn't be able to conclusively prove the idea fixes the issue in larger models, if you know what you are looking for you should be able to validate that the method works in general down to very small models.

That said, consumer grade cards should be able to train an 8B model with quantization, so you might as well train the whole thing.

It doesn't need to be two huge models. If there is an advantage to doing this, I'd expect that you would see it even in a small test case. I'm sure we'll see something by the end of the week if not earlier if there's something to it.

The important and popular ones are absolutely available, but those are usually important because they have entered the realm of "common knowledge," at least in a particular sub-field. These are going to be at the top of the list when it comes to digitizing useful historic records. It's fairly easy to OCR a PDF, so as long as someone with some time decided "hey, this might be useful" then you'll probably be able to find it.

If you're doing databases then you've almost certainly been exposed to Codd's work, if not through his papers and books, then at least through textbooks and lectures. There are countless blogs, lecture series, and presentations that will happily direct you there.

The challenge is that there's also a mountain of work that never really got much popularity for whatever reason. Say a paper was ahead of it's time, or was released with bad timing, or simply kept the most interesting parts until the end where few people might have noticed. It's these sort of gems that are hard to find. It's hard to even know how many of these there are, because they are by definition not popular enough for most people to know about them.

I think this problem comes down to two core issues: discoverability and terminology.

You're going to be lucky if a paper from the 70s or 80s is available in a searchable database at all. That means someone bothered to scan it in, and OCR it since then. Even for the few papers that are searchable, they are old enough that they probably won't catch anyone's eye unless they are desperate.

Of course then there's also the problem of knowing what to search for. Programmers love to invent, reinvent, and re-reinvent terminology. It's only gotten worse with every other developer running a blog trying to explain complex ideas in simple terms.

The entire field of ML is a perfect example of this. I remember talking to my father about all sorts of new developments in ML back in the early 2010s, and I was quite surprised when he told me that he learned a lot of the things I was talking about back in the 80s just named a bit differently.

In most cases it ends up being a question of how much time you can put into any given problem. If I spend two weeks to find a paper that would have taken me a week to reinvent, then am I really ahead? If the knowledge wasn't important to enough make it into textbooks/classes/common knowledge then attempting to find it is akin to searching for a particular needle in a pile of needles.

That really depends on quite a few other factors: how big is the team? What development methodology do they use? Does the leadership understand how to manage and direct a rewrite? Are there people that understand the full scale and scope of the system? Does the system interact with legacy components that can't be modified? Are there political factors in play? These are just a few of the questions that can change the outcome of any given rewrite.

You mentioned hidden bugs, but what about hidden "features" that may be a critical part of existing business processes for core parts of the company? Developers really like to believe they are at the center of the wheel due to the complex work they do, but a lot of the time they are not the ones that actually create the cash-flow.

I've been part of rewrites that have succeeded tremendously, but I've also been privy to utter failures that have cost millions, and led to entire teams getting sacked.

10 years isn't really all that much, is it? From my experience that's around how long it takes for developers to get a big head about how much they know, but 5 years less than what it takes to learn to respect how much they actually don't know about the different aspects of the field, and the real scale of challenges that have to be solved (both the technical, and the human).

Also, not all experience is equal. Someone that's spent 10 years working on 4 or 5 different systems in totally different problem domains, written in totally different languages, and operating in totally different ecosystems is going to have a very different view of development from someone that's spent 10 years doing essentially the same thing over and over again.

This guy seems to have a very focused view of the correct approach to problems. He's familiar with the tools that linux offers (which I agree are great), but he doesn't seem to respect the scale of specialization it takes to use and maintain those tools effectively on a large scale. Also, there is no mention of the cost to rebuild existing systems in terms of developer time, the mental cost to re-train all of the developers, as well as the time to migrate and train the users.

Ironically, I remember getting into debates like this back in the mid-2000s when I was first starting to think I had it all figured out. The points I made back then were more or less the same things I see now in the article above. It's quite nostalgic, though it definitely makes me feel older than I like.

Just looking around, general available figures for public internet (as opposed to tor) suggest that anywhere between 0.1% to 1.0% of users have JS disabled. These numbers have also been consistently going down over time. That's a fairly small number to dictate how a system should be designed.

I'd also argue that if tension 1 were really a problem (i.e., Reddit staff were wrong), Reddit would be obviously going downhill, while tension 2 can fester as organizational debt for years before exploding, if everyone is well-intentioned.

I have found the issue to be not so much a matter of the two factors that you've outlined, but more a consistent downward trend of the admin staff, towards a stronger disconnect with the community. There has been less communication, and the communication that has happened has been less clear and less consistent. Even in this entire drama, reddit's response has come through a single point of contact.

For a site like reddit to work, the administrators really need to be able to also participate in the community at large. They need to have firm, definite rules and guidelines of what they will and will not do, and how they will or will not help. They need to make themselves available to the volunteer staff that help run these numerous communities.

This I think is the root cause of both of these tensions. The community simply doesn't know what to expect from the admins anymore.

I took a look at the website you linked at the bottom of your post.

The first thing I noticed is that most of the front-page articles are testimonials. Fortunately there was an article written by the author right at the start, however as soon as I started reading I noticed another problem. Almost every reference is to other articles on that same site, which themselves link deeper into the site. Occasionally I'd hit some articles from psychology today (psychology magazine, not a journal), and other popular media sources. I did not find any references to proper scholarly articles though.

A quick search of Google scholar turns up no articles to back up most of the major assertions he makes on here. Now that's certainly not enough to dismiss the site outright, but it's certainly enough to make me question what exactly he found, and how well he is interpreting the existing results. I have no trouble accepting the existence of porn addiction, but I do believe that making a case as strong as the one you seem to be making requires much stronger evidence than what you have presented.

> Civil eng is very conservative in terms of the kinds of language and graphics that can be used to express a design.

I would argue that programming is far more specific in terms of the kind of language can be used too. In fact each such language tends to be described in exhaustive specs.

> Anyone doing the equivalent of currying or macros (making up ones own language) would be thrown out. I would think its probable that when programming is as old as engineering its modes of expression will be similarly limited/standardised.

Both things are very broad when it comes to what can be made using those languages. A civil engineer may use his language to build a house, a sky-rise, and a nuclear power plant. Each of those will have different complexities, and a different requirements of knowledge and qualifications. In fact I imagine the Engineer working on the latter will know how to do a lot of things that the Engineer who works on the former would consider to be akin in complexity to currying and macros.

The situation is the same in programming. Some people may be working on projects currying, macros, and other techniques are a major benefit. These are after all extremely powerful tools. Just like with the Civil Engineer, the challenge is knowing how to use them properly.

> shouldn't it be a commonplace thing that doesn't take so much work to get around to explaining and using?

Why? It's a reality of the world that more complex things take more time and more effort to learn. However, often that is because these more complex things allow you to do a lot of very useful things much more efficiently. I would prefer to drive over a bridge built by an engineer who learned all those difficult equations, material properties, and buildings codes as opposed to a high school kid with a few physics courses under his belt.

Programming is similar in some effects. As you get better and better you acquire more and more tools to do what needs to be done. Now granted, if you are working on an interface that needs to be easily accessible to the widest range of people it makes sense to simplify. However, cleverness has it's place in code that is expected to be read by specialists.

In the end, even if you avoid all the clever tricks and shortcuts you know, a large enough project will still be utterly inaccessible to a novice. The real challenge of projects that complex becomes less about the specific detail of how a piece works, but more about how all the pieces work together. If you're skilled enough to follow the design of a project like that, I don't think it's too much to ask that you either know these "clever" techniques, or you should be willing to learn.

Looking at your code you linked in the article, I think part of the problem is the fact that there are entire pages of code without a single inline comment. When you're doing these clever things you really need to document every logical step in order to understand and verify your through process later on. You also have to be ready to accept that sometimes you will mess up in your cleverness. In fact, If you are getting a lot edge cases that's a good signal to go back, re-read your comments/design notes, and find where you could improve your approach.

Ironically, I would argue that go channels are actually an example of doing something "clever" the correct way. These channels are very effective at separating a single concept from a whole pile of abstractions, and doing a lot of clever interactions beneath the hood in order to ensure it's all effectively synchronized. In other words, using go channels is using the same type of "clever" techniques once they've been abstracted away.

I feel a major problem with mathematical proofs is that the language of these proofs is so complex that in order to understand what it even says you have to be extremely smart, and willing to put in a huge amount of time in the first place.

A proof is in effect a "program" that describes the logical conditions of whatever it is trying to show. Invalid proofs are those with "bugs". Perhaps if the language of math was a bit clearer, it would not be such a difficult to understand field.

I wouldn't say he's describing "adding another system to the mix," at least not in a traditional sense; it's more along the lines of two distinct systems working towards a similar goal.

What he described is an auditing system with some particular policies of interest to a specific use case. Such a system should not have any direct access to the main system, and should ideally live in a fully segregated environment with tightly controlled read-only access to a copy of the data being audited.

The whole idea is that this system would not announce its presence on the network in any way so that the attacker is more likely to miss it. Even if the attacker does know that it's present, he should not know all the checks and validations that such a system uses to detect suspicious behaviour. Hell, you could air-gap the entire thing and just copy over data dumps by using USB sticks.

Granted, even in that situation you could get something like Stuxnet which may compromise the machines. However, if you have the resources to build another Stuxnet, chances are you don't really need to get into a bank network.

While that is a perfectly reasonable solution to the minimal example I provided, there may be reasons to abstract away the concept.

Personally, I like having some sort of structure in memory that allows me to access and query contextual information. It's amazingly useful if you are trying to link a variety of different elements. In fact it's practically unavoidable if you are writing highly dynamic code. Trying to manage this sort of system with a few variables and functions would become a nightmare.

In fact, those dynamic situations are where comments describing layout are absolutely critical. In these situations you may be using your class as a generalization, and tracking the full extent of this sort of interaction could take a ridiculous amount of time.

I think a lot of people use the "self-documenting" excuse without knowing what "self-documenting" really is, and what it brings to the table. Self-documenting code is a good fit for a reasonably small and straight forward function. If you have a function called get_thingy() in the context of some sort of "thingy container" then you probably don't need to go out of your way to explain "This method gets a thingy from the thingy container."

This can even be expanded to more complex single-purpose functions that are not exposed directly to any external APIs. For instance, if you have a functioned called restart_thread() which first tries to stop a thread, then stats a new one then you might be safe.

However, as soon as you leave the realm of the obvious, comments are a must. Clearly in situations where you're trying something clever, like that Q_rsqrt function, missing comments will boggle all but the most specialized professionals.

Unfortunately people often forget about another important type of comments; those that explain what function this piece of code serves in the program. I've lost count of the number of "Context"s or "Interface"s or "get_data"s I've seen when trying to read someone's code. In those situations a single line like "This context links the rendering pipeline to the physics simulation" would go a very long way.

Instead when I see this sort of code it will either not have any comments at all, or it will be a dissertation that all the theoretical uses of the piece of code (Omitting what it's actually used for in the current context).

This has always been a pet peeve of mine with projects like rails. Usually when I work on a piece of the system I'll be working on a single component, be it user, product, report or whatever. That means I will either have a bunch of different directory trees open in my IDE (Which can be bigger than a screen for large projects), or I will be using some quick-open tool (Which still is great when you know what you want, but less so when you just want to get a quick overview of a component).

God forbid I try to edit multiple components at a time, and forget browsing my repo on github. Organization like this would make life so much easier in so many ways.

05:21 -!- mode/#linode [+b !ryan@54.228.197.*] by akerl

05:24 -!- ryan| [~violator@37.235.49.168] has joined #linode

05:24 < ryan|> quite rude of you

05:25 -!- ryan| was kicked from #linode by akerl [ryan|]

05:27 -!- root__ [~h@vmx13318.hosting24.com.au] has joined #linode

05:27 -!- root__ is now known as ryan||

05:27 < ryan||> Quite rude out of you

05:27 < ryan||> To ban me like that

Really puts into perspective the difference in the levels of skill involved.

When dealing with someone of this level, they really should have just notified everyone immediately. There's no telling what info these people have now.

Having multiple modes of work is critical in the field. It's an unfortunate reality that the theoretical complexity of programming projects is limitless. Unfortunately that means you simply can not have sufficient resources to do things correctly. This is where the balance of not just approaches, but also development methodologies comes in to play.

Whenever I'm working on my personal project I do things that many modern developers would balk at. I'll design and redesign components left and right, I'll leave components in a half completed state, and I'll write the bare minimum of tests to ensure the system works. This is the risk that comes from approaching programming as an art-form. However, as any other artist, I will not go out there to show off my project until it's done, so I don't have to worry too much about pissing people off with my erratic style and questionable intermediate decisions.

On the other hand programming as a profession must be wholly different. If I am getting money for code the expectation is that people want to see whatever they're paying money for. In this case I don't have the opportunity to screw around and try to match the perfect set of ideas to the problem set. Instead I will give myself two or three iterations to do what I've been paid to do, and then I'll clean it up so it satisfies the technical requirements.

Fortunately the latter style of programming tends to lend itself a lot better to TDD, agile, responsiveness and all those other terms we love to throw around until our clients give us that blank stare. At the very least projects you get paid for should have some sort of spec, so you'll be able to say for certain whether you have or have not done what's expected of you.

I'm not suggesting they should have broken up all the teams and made new ones. There are much smarter ways to merge teams that involve gradually easing them together. However, having five teams in the company doing nearly identical things is not "gradually achieving" anything.

If (big if) teams had development plans they had to follow, then those plans should have been adjusted so that eventually all these teams were working towards a common purpose. If you just leave those teams alone and hope for the best not only are they going to avoid any chances to work together, but they will often go out of their way to ensure they don't happen.

From my perspective, AMD really screwed up the AMD/ATI merger in the worst way possible. When AMD bought ATI they were well positioned to beat everyone to market to release their APU chip. However, the problems started almost immediately when neither AMD nor ATI did anything to combine resources. They should have put in the time and money to merge teams at all levels of the company. Instead they built a few small groups that included senior personnel from both units, then they left most of the lower level teams to work on whatever they were working on previously. Never mind the fact that there were a lot of really clever people all over the place ready to contribute really great ideas.

In other words, there was zero global direction down the ranks. It was just business as usual; keep doing what you've always been doing, and maybe we'll show you some nice slides a few times a year about how great APUs will be. This lack of organization meant that no one had any idea what anyone else was doing. Worse-- even if you wanted to find out there was absolutely no company-wide documentation or organization on anything. Your only hope for getting information was hoping one of your co-workers had bookmarked some magical page with the info you required. l This just got worse when you accounted for the problem of elitism. The hardware teams were just so much better than the software teams. After all, software is easy, so what sort of useful input could those code monkeys offer. And far be it from the software teams to actually talk to someone from the QA teams; those QA people were beneath notice. Finally, add in a very wide distribution of personnel seniority, insane levels of paranoia about job security, grade-school level office politics, and completely disparate management styles, then hit blend.

So really, the results are not at all surprising. You can't have two companies pretend to be one while playing tug-of-war, and still be competitive.

If your memory usage and gc speed are of critical concern, then you should really know that ruby is just not going to provide the tools you need to handle that. That's like using a hammer when you need a screwdriver. Even the most modifiable hammer is still meant for pounding, not screwing.

Whenever I see "I'll write in Ruby and optimize in C when I need to" I assume the context of "my program is well suited to be written in Ruby." For instance, I certainly wouldn't write embedded code in ruby. However, I would venture to say that these situations are the exception rather than the rule.

These days most software has access to fairly reasonable computational resources. If you're writting microcode for some industrial control system, or calculating stock price variations with micro-second resolution then certainly stick to C or ASM or what have you. However, in an age when even a fridge will have a few hundred MB of memory, and when a lot of code is meant to be just "good enough" I think even ruby will suffice.

I would venture to say the triviality of writing hybrid code is a function of how often you have done something similar, and how well you understand the underlying principles. This is really the case with any software problem.

Ruby provides some very useful features, but a lot of those come at rather high costs. Those costs are exacerbated if you do not understand how the internal components of the language are laid out, and the purpose behind this layout.

Ruby tries to be a lot of things to a lot of people, so then when people learn how to use it for one task they automatically assume that their skills will carry over. This sort of approach might work reasonably well with a straight-forward compiled language, but this it simply can't be that easy for an interpreted language like ruby, with it's reams of special context sensitive behaviour.

For example, consider the "pointers to pointers to pointers" complaint. Nothing is stopping you from having a low level C struct for all your performance sensitive data. Granted, you would have to write a ruby wrapper for this data for when you need to export it back out to ruby, but wrapper generation could be automated.

Sure, you could just say, "You know what. I'm just going to use C" but what if your project could really use some of those ruby features outside of the performance sensitive parts of the code? It's always a tradeoff.

From Wikipedia: "There are over 67,000 subreddits to peruse, with the default set being (as of October 18, 2011[4])." I could not find a specific number for the admin staff, but given the fact that wiki lists 11 staff members for reddit that number can not be particularly high; in fact the reddit admins steam group has 6 members. I sincerely doubt that anyone is capable of filtering 10,000 subreddits every day in search of offensive material.

Finally, even if somehow reddit manages to completely and utterly block all things they deem to be CP related, these people will just move to another, probably harder to track venue. All in all, this entire move is rather pointless action in response to people that want the feelgood sensation of "Protecting the Children."

Casual gamers are driven by trends and early adopters. A lot of these people know what Kinect is because their "nerd friend Alphonso" won't stop talking about how cool it is. Don't knock the early adopter crowd just because they are small; while they may be that, their extended social circle is generally large enough to sway even the mainstream numbers.

Correct me if I'm wrong, but I was under the impression that when most people think of a DS, or any other Nintendo product, the Nintendo properties are the first thing on their minds. If you are getting a DS you probably want to play Mario, Zelda, Pokemon, or Metroid. If those do not interest you, then you are still more likely to be interested in some sort of JRPG or maybe a fighting game.

I am a gamer, and I was not even aware that there was any sort of market for LEGO Harry Potter or Sims 3 on a DS. In this case I think the stats support my view: http://www.gamestats.com/index/gpm/nintendo-ds.html

You may note there are only four non-Nintendo properties in the top 20 popular games, and of them only one (GTA) has a native iOS port, while one more (FF6) could technically run on a SNES emulator. I hope no one asking you for game platform advice was hoping to play any of the most popular portable games of the day.