HN user

lignuist

1,394 karma
Posts26
Comments339
View on HN
www.reddit.com 12y ago

Anything related to Tesla has been secretly banned from /r/Technology.

lignuist
8pts2
googletranslate.blogspot.com 12y ago

Google Translate - now in 80 languages

lignuist
1pts0
rt.com 12y ago

Indian minister in charge of secure email policy uses Hotmail

lignuist
2pts0
www.bbc.co.uk 12y ago

Japan approves new state secrecy bill to combat leaks

lignuist
2pts0
www.geek.com 12y ago

Swear in a private Xbox One Skype call, get banned from Xbox Live

lignuist
4pts0
www.japantimes.co.jp 12y ago

NSA asked Japan to tap regionwide fiber-optic cables in 2011

lignuist
5pts2
www.bloomberg.com 13y ago

Hong Kong Banking System Outlook Cut by Moody’s on China Risks

lignuist
29pts15
www.guardian.co.uk 13y ago

MI5 feared GCHQ went 'too far' over phone and Internet monitoring

lignuist
20pts1
www.youtube.com 13y ago

Buddy Cup: When two people toast, they become friends on Facebook.

lignuist
1pts0
mtgoxlive.com 13y ago

Bitcoin hits 250$

lignuist
8pts5
vimeo.com 13y ago

Hacking gravity: BALANCE

lignuist
3pts0
www.edmontonjournal.com 13y ago

Ten of the biggest product flops of all time

lignuist
2pts0
www.microflight.com 13y ago

The Worlds Smallest R/C Blimp - Nanoblimp

lignuist
2pts0
www.scanomat.com 13y ago

TopBrewer - Coffee Tap

lignuist
1pts0
edition.cnn.com 13y ago

The Internet is a surveillance state

lignuist
318pts174
www.youtube.com 13y ago

New York Day: Time-Lapse Video

lignuist
3pts0
gardening.stackexchange.com 13y ago

How can I tell if a plant given to me is patented?

lignuist
104pts60
en.wikipedia.org 13y ago

Wikipedia: Comparison of feed aggregators

lignuist
2pts0
metro.co.uk 13y ago

Danish TV used Assassin’s Creed screenshot in report on Syria war

lignuist
4pts0
christianengstrom.wordpress.com 13y ago

European Parliament censors citizens trying to contact MEPs

lignuist
2pts0
bitcointalk.org 13y ago

Bitcoin: Soft block size limit reached

lignuist
49pts59
www.h-online.com 13y ago

Frosty attack on Android encryption

lignuist
24pts15
flock.codeweavers.com 13y ago

CrossOver for free today.

lignuist
16pts0
mdn.mainichi.jp 14y ago

Another reactor to shut down, leaving only 2 units online in Japan

lignuist
1pts0
blog.ioactive.com 14y ago

SSL Traffic Analysis on Google Maps

lignuist
15pts0
www.dacia.de 14y ago

Car vendor fakes hacked website.

lignuist
1pts0

Nitpicking much?

As I wrote above, by making sure that I use a placeholder that does not appear in the data, I make sure that it does not cause the issues you describe. And if I was wrong with that assumption, I can at least minimize the effect by choosing a very unlikely sequence as placeholder.

I really see no issue here. How do you find valid grammars for fuzzy data in practice?

I'm not sure if interchangeable is the right word. 'phởne' and 'phone' yield different result lists.

At least Google seems to have a way to detect visually similar letters.

New Haxe website 12 years ago

The idea is to be able to write code in one language and use this code in many different languages.

I used that strategy for parsing gigabytes of CSVs containing arbitrary natural language from the web - try to get these files fixed, or figure out a grammar for gigabytes of fuzzy data...

My approach never failed for me, so telling me that my strategy does not work is a strong claim, where it reliably did the job for me.

Your examples are all valid, but what you are describing are theoretical attacks on the method, while the method works in almost all cases in practice. We are talking about two different viewpoints: dealing with large amounts of messy data on one hand and parser theory in an ideal cosmos on the other hand.

What if there is #COMMA, in one of the fields (but no #COMMA#)?

What should happen? Since #COMMA is not #COMMA#, it gets not replaced, because it does not match.

Please keep in mind, that I replied to suni's very specific question and did not try to start a discussion about general parser theory. In practice, we find a lot of files that do not respect the grammar, but still need to find a way to make the data accessible.

You just choose a placeholder that does not appear in the data. You could even implement it in a way that a placeholder is automatically selected upfront that does not appear in the data.

When it comes to parsing, the thing is that you usually have to make some assumptions about the document structure.

I was referencing to "What if the character separating fields is not a comma?".

And there it clearly works. I used this technique a few times with success. If you find a CSV file that has mixed field separator types, then you probably found a broken CSV file.

You can replace all commas with a placeholder (e.g. "#COMMA#"), replace the delimiter with a comma, parse the document and then replace all placeholders in the data with ",".

My perspective on art is a reaction on the elitism of the art scene, so basically my comments are art.

Edit/addition: Honestly, I could have much more respect for this project, if Wu-Tang made it only accessible to homeless people, or only to prisoners, but effectively, they make it only accessible to the riches. I really do like the Wu-Tang Clan, but I am really not impressed by this stunt.

He continued: "I don't know how to measure it, but it gives us an idea that what we're doing is being understood by some. And there are some good peers of mine also, who are very high-ranking in the film business and the music business, sending me a lot of good will. It's been real positive.

So Wu-Tang Clan fans in Kazakhstan or Tanzania (or even every country other than the U.S.) will probably never be able to listen to this album...? I guess these will be the people who don't "understand" what Wu-Tang Clan is doing, while only the privileged ones "understand" the concept.

That's artificial shortage, not art (not talking about the music itself).

Pending Comments 12 years ago

This is what I feel too. It expect it to streamline the comments and kill the discussions.

Is Microsoft circa 2014 worse than Google, Apple, or Facebook? We're not nearly as organized as we'd need to be to be as evil as you might think we are.

Microsoft is not any worse than the other companies. They are all at the same terrible level.

But Microsoft became a bit better over the last years, I would say.

Citing the AngularJS FAQs:

Does Angular use the jQuery library?

Yes, Angular can use jQuery if it's present in your app when the application is being bootstrapped. If jQuery is not present in your script path, Angular falls back to its own implementation of the subset of jQuery that we call jQLite.

Due to a change to use on()/off() rather than bind()/unbind(), Angular 1.2 only operates with jQuery 1.7.1 or above.

http://docs.angularjs.org/misc/faq

Human Corrected Translations for 1 cent per word

"Per word" of the source language, or the target language? Sum of both? What about languages which have a different concept of "words" in written text (e.g. Chinese, Turkish, ...).

And by the way... "cent" of which currency? :)

Edit: I just saw that the list of supported languages does not contain languages with "exotic" types of word boundaries (yet).

I personally think it's a bad idea to mix sport and politics. We already had it in 1980 and 1984

As soon as there is big money involved (which is definitely the case with the Olympic Games), sports and politics are always mixed. Countries have to spend billions in order to make the games happen. This is not just a few people doing sports for just the sports.

I think putting in Airplane Mode should be a good idea.

And then you still would have to trust this Airplane Mode.

Honestly, I have absolutely no idea what is happening in my phone. Is it maybe still collecting data while it is in Airplane Mode and sending it somewhere once it is set back to regular mode? Or is it even sending data all the time, because the NSA knows that never a single plane has crashed because of active cellphones? Probably some smart people out there are checking their phone's internals and activities more than I do...