NLTK is great for _learning_ NLP, but Python is much too slow for scalable deep NLP (by which I mean tagging and parsing, as opposed to TF-IDF etc). Also parallelization can become a problem because of the GIL. It's a real shame they chose Python actually, because otherwise it's a superbly structured, documented, and maintained project.
HN user
gilesc
Graduate student, bioinformatics, Univ. of Oklahoma Health Sciences Center Machine learning, text mining, gene expression studies
It's largely been supplanted by Ruby?
Like sports, instruments, and any other developed skill, there will be days you feel like it and days you don't. The key IMO is to pace yourself: don't overwork on days you feel like it, and force yourself to work -- even if just a little bit -- on days you don't.
For software, having your code in a public repository like GitHub provides some socially-based motivation to keep your projects active. Just the simple act of regularly committing small changes can provide you with a sense of momentum -- and bonus points for raising (and fixing) issues, etc.
For days that you aren't on your A-game, it's also helpful to have made a TODO list from a day when you were thinking more clearly, so you can work on a relatively simple task just to maintain momentum.
Another key is to have a "big picture" goal that your projects are helping you towards. There's no reason you can't start now putting together the basic structural code (say, some core machine learning algorithms) for a later startup -- or even try your hand at writing an end-to-end web app and hosting it for free on Amazon. Whatever your end goals are, you'll be more motivated if you are writing code that helps you get there, not just code for learning's sake.
Maybe referring to this:
While they were still choosing targets that either "deserved" it for some moderately reasonable ideological reason (Sony) or should have top-notch security (FBI), yes. Now that they're choosing targets at random, no.
Especially since this is all so likely to end in an acceleration of government crackdown on web freedoms.
Interesting that the author of the fake in-browser Finder apparently uses Dropbox.
Reads like a HOWTO for burnout.
Working on this, if anyone wants a corpus (positive examples from twssstories.com, negative from fmylife.com)
In the US, PhD students in biology are paid tuition plus salary. So assuming his undergraduate student loans were subsidized, a PhD who decided to teach high school would not be in a much different financial position than a recently graduated Science B.Ed.
Not to mention that there are magnet high schools that do very much like the perceived prestige of having PhD level instructors. Your typical public high school wouldn't care, of course.
Well, first of all, a traditional student will be out by age 26-27 not 40. Secondly, PhD students in the sciences are paid. Not much, admittedly -- about $22k, on average -- but enough that you have positive cashflow.
Well, bioinformaticians are far from uniform in terms of language usage (my mentor staunchly defends his use of VB6, for instance...), but this link shows people using Python for structural bioinformatics as early as 2001: http://mgl.scripps.edu/people/sanner/html/talks/PSB2001talk....
To be fair, there are also private research foundations and private sector companies -- i.e., drug or medical device companies that also hire biologists. You can also pivot somewhat easily into public health and work for e.g., the Department of Health. And if all else fails, teach high school or lower biology.
Ha! Thankfully, Perl seems to be mostly going out of favor in bioinformatics in favor of a Python / R combo.
I think this article is a bit sensationalist. Yes, there are niche techniques in biology, but many -- Western blots, qPCR, cell culture, transfection -- are transferable to almost any other biology lab.
The article's core advice -- make sure to acquire transferable skills -- is certainly a good idea, though.
As a bioinformatics PhD student, life is great: I get to learn those programming and math skills specifically mentioned by the article as being transferable, along with some of the other more esoteric skills. (Moral of the story: choose bioinformatics! :)
This looks like a slightly better cake. After all, cake provides a persistent VM and you can run scripts with cake with !#/usr/bin/env cake.
It's nice that it doesn't require ruby and has a few additional features like search & install from clojars, and doc search. But it probably doesn't justify the cost of switching since I can get those functionalities from cljr.
What would really be nice -- the maintainers of lein, cljr, cake, and now jark all need to get together and work together instead of providing competing and slightly different tools.
Although I also dislike .NET, it seems the primary complaint here about .NET is that it abstracts programmers too much away from the bare metal. But doesn't that apply equally to non-MS languages like Python and Ruby? In fact, isn't abstraction in general a good thing?
The author also bewails .NET's lack of configurability. But what is it about C# that instantly defiles any programmer who touches it, that doesn't also apply to, say, Java?
see also: https://github.com/weavejester/clucy
It's a great idea technically, but I'm not sure whether you can sell a social networking app. Social networks are valuable in proportion to the number of users, and pay-for-access adds a huge barrier to growth.
They might be better off running some unobtrusive ads.
Independence / self-direction. Searches google / manuals for answers instead of asking you every few seconds.
This is absolutely gorgeous. The only thing the standard RGui has on it is basic emacs/bash-ish keybindings (C-e, C-a, etc.) for navigation.
The best toolkits are probably in Java:
-Stanford's Tagger, Parser, and NLP Core
-Apache OpenNLP
-Lingpipe
Many smaller components are made to be compatible with IBM UIMA (of Watson fame), so they are able to be integrated into a pipeline somewhat easily. For examples of this in biomedical TM, see http://u-compare.org/ .
People will kill me for saying this, but truly: Python's performance isn't adequate for large-scale text mining, _especially_ if you want to do deep/full parsing. Shallow parsing as shown in this package's demo is more feasible.
I personally find NLTK convoluted, but in its favor, it does have readers for a TON of corpora, which is really nice.
Oversimplifying things a bit, these writers are saying that scientific objectivity is impossible. This line of argument tends to reduce the credibility of science in the eyes of the public, with negative consequences for research funding. This is a bad thing because, even if not "objective", science has singlehandedly transformed life for the better in the past century by reducing disease, poverty, manual labor, etc, and should continue to be funded.
I like and use it, but it's based on reverse-engineering Pandora's encryption keys, so every once in awhile (when Pandora changes the keys), it breaks the client, which is annoying.
If she ends up losing the data permanently, that will severely hamper her ability to apply for future grants (which usually require preliminary data), and lose funding. This offense is definitely self-punishing.
I'm a student at OUHSC, and her husband, Dr. Janknecht, is one of my professors. Although failing to make a backup is obviously stupid, he is an extremely competent researcher (I don't know her). Clearly, calling the data "a cure" is a bit of hype, but that exaggeration might help in convincing the thief to bring the laptop back.
About computer knowledge in biological research, though -- the state of things is generally abysmal. The average biology Ph.D. can use Excel to find means, SDs, and do t-tests, and that's about it. Even my boss, who specializes in bioinformatics, still uses VB6+MSAccess shudder. Most probably don't know that hard drives CAN fail.
Yet, researchers are fiercely independent and would definitely resist any heavy-handed mandates from campus IT forcing specific OSes or regular backups.
Interesting choice of metaphor given the original topic...
I'm a biochemistry grad student, and my school is just now considering offering a (bio)statistics course for the first time... But parent poster is right, chi-squared is usually as complex as it gets.
There is also a sub-group for HNers who are also students: http://www.facebook.com/home.php?sk=group_152776264761584...
Hence "Yes, Mike, I Have Stopped Beating My Wife".
A phone could give FB plenty of new chances to capture data. What if, instead of sending a text message to someone, the FB phone sends a FB message or initiates a FB live chat session with them instead. If the FB phone became ubiquitious enough, they could challenge the space currently owned by texting (which would be really annoying for developers and the market, but great for FB). FB locks everyone in, AND they get to use all that social data too.