IQVIA is a Global CRO with like 80-90k employees. They do two things really well. Running clinical trials (Phases I-IV) and vacuuming up healthcare data (claims, EMR, sales, etc., https://www.iqvia.com/insights/the-iqvia-institute/available...)
HN user
kprybol
Counterpoint to that is when a recruiter initiates contact they are often very coy until they have you on the phone, frequently not even sending a JD until after that first conversation (where this line of questioning usually arises). I get wanting to provide an engaging response but hard to do so until the recruiter has provided sufficient information.
Certilytics | Hiring Data Engineers, Product Managers, Analysts, Developers, and more | Remote: US Only | Full-time
Certilytics provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights. Our team represents a dynamic infusion of multidiscipline, which includes actuarial, data, and behavioral scientists, IT engineers, software developers, nurse clinicians, and experts in public health and the health insurance industry. Certilytics has extensive experience working with a diverse set of customers, including large self-insured employers, health plans, pharmacy benefit managers, government programs, care management companies, and health systems. These relationships with various data providers and customers allow for rapid data ingestion, validation, and enrichment, as well as streamlined delivery of analytic dashboards, outputs, and visualizations to our customers. Our unique approach allows for developing the most accurate financial, clinical and behavioral models in the industry.
Why Certilytics! Access to one of the most extensive clinical datasets in the industry that includes medical claims, pharmacy claims, and laboratory data. Impactful work. We're big enough to have the freedom to take on interesting projects but small enough that your work is always important and highly visible within the organization. Remote friendly. The Certilytics team is distributed throughout the US and has regular in-person working sessions for all of those little things that are hard to accomplish over Teams.
See our open jobs: https://jobs.silkroad.com/CarewiseHealth/Certilytics
Certilytics | Hiring Data Scientists and Data Engineers | Remote: US Only | Full-time
Certilytics provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights. Our team represents a dynamic infusion of multidiscipline, which includes actuarial, data, and behavioral scientists, IT engineers, software developers, nurse clinicians, and experts in public health and the health insurance industry. Certilytics has extensive experience working with a diverse set of customers, including large self-insured employers, health plans, pharmacy benefit managers, government programs, care management companies, and health systems. These relationships with various data providers and customers allow for rapid data ingestion, validation, and enrichment, as well as streamlined delivery of analytic dashboards, outputs, and visualizations to our customers. Our unique approach allows for developing the most accurate financial, clinical and behavioral models in the industry.
Why Certilytics!
Access to one of the most extensive clinical datasets in the industry that includes medical claims, pharmacy claims, and laboratory data.
Impactful work. We're big enough to have the freedom to take on interesting projects but small enough that your work is always important and highly visible within the organization.
Remote friendly. The Certilytics data science team is distributed throughout the US and has regular in-person working sessions for all of those little things that are hard to accomplish over Teams.
Tech Stack: Python, TensorFlow, Kubernetes, AWS, Spark, Scala
See our open jobs: https://www.certilytics.com/contact-us/#careers
Certilytics | Louisville, KY | REMOTE (US Only) | Machine Learning Research Engineer
Do you enjoy reading the latest machine learning research on Arxiv? Do you challenge yourself to reverse engineer interesting papers? Do you seek to apply existing algorithms to new domains and develop creative and novel solutions to difficult problems? If so, come join our team at Certilytics!
Certilytics, Inc. provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights.
As a machine learning research engineer, you'll be responsible for designing and running experiments to bring the latest deep learning advances from the literature to our products. As part of the data science team, you will be responsible for building models for clinical and financial risk prediction, performing original research, and contributing to a proprietary machine learning library. The ideal candidate will have a strong background in natural language processing and familiarity with the inner workings of RNN’s and transformer networks (Join a flexible, energetic team in bringing the best of deep learning to healthcare.
Apply here: https://jobs.silkroad.com/CarewiseHealth/Certilytics/jobs/20...
Does this do experiment tracking similar to MLflow? I’m trying to figure where these two overlap and where they diverge.
A subscription to safaribooksonline.com ($399/year) is my favorite way to spend part of my learning budget. Having access to the entire catalog of O’Reilly (and it’s affiliates) books is awesome. Access to conference recordings from Strata is really nice too.
Basically just the ability to run on Linux.
Any chance of this being offered for on-prem install in the future? Looks interesting but cloud only makes it a no go for my team.
From my experiences (currently work with several Fortune 100 health insurers/benefits managers, and have previously worked for another large insurer, a major academic medical center, and a large pharma company), healthcare organizations tend to be rather cloud adverse (most of our contracts very explicitly forbid us from using any form of 3rd party cloud computing). So while I agree that much of the heavy lifting will shift to the cloud (or already has), I expect health analytics will continue to favor on-premises solutions (GPU’s still tend to be pretty rare compared to CPU based clusters but are slowly becoming more common).
That's fair. Google's recent paper on predicting patient deaths is another good example of this (logistic regression + good feature engineering performed just as well as their deep learning models, and the logistic regression has the added benefit of being significantly more interpretable and as a result, actionable).
It'll be interesting to see when specialized ML focused silicon will become readily available. Right now I find ML libraries that are able to run on blended architectures (any combination of CPU and GPU's) much more exciting/impactful than TPU's. The ability to deploy on just about any cluster a customer may have available is huge.
I've always viewed DeepMind as more of a skunk works program and less as a profit driven enterprise. DeepMind exists primarily to push the limits of what can be done when you put group of leading researchers together in a room, provide them with nearly limitless resources, and simply tell them to "go". I expect some of that effort to eventually trickle down into Google's consumer products (maybe a healthcare focused version of AutoML https://cloud.google.com/automl/). Google has already done a lot of work on the HIPPA side of things (https://cloud.google.com/security/compliance/hipaa/)
Being rules based isn't necessarily a bad thing or disingenuous. I develop healthcare AI products (ML/DL researcher) and we actually aim to be able to translate our models into a rules based engine (find a strong signal, interpret/understand model well enough to translate/embed into a rules engine, look for a new signal in our models, rinse + repeat). We end up deploying a mix of rules based and true ML based models into production but it may not be immediately obvious to the end user which type of model they are using.
Healthcare data can't be shared the way the browser histories, cell phone location data, etc. can. It's a completely different set of a rules that people have to play by (HIPPA for example). I build machine learning systems using healthcare claims and EHR data and without the direct cooperation of several large insurance companies (and access to their data) we'd be dead in the water. Even with access to their data there are incredibly strict limits to what we can and can't do with it. You can't just go out and collect healthcare data the way you can many other types of data.
It depends on the use case. Our work primarily revolves around extending Spark with custom pipelines, models, ensembles, etc. to be deployed into our production systems (petabyte scale). Scala was really the only way to go for us.
Kafka too
Scala is used pretty heavily in the big data world, particularly if you are working with Spark.
Not sure how easy it is to find outside of boutique pets stores (just happen to live near a fantastic one that delivers for free) but I highly recommend the Orijen brand. Have tried numerous other “premium” brands with my dogs over the years but no matter which brand I was using, I always managed to run into at least one vet than had less than stellar things to say about brand x, y, or z. Maybe it’s just coincidence but in the 7+ years I’ve been feeding my dogs Orijen I’ve yet to encounter a vet that anything other than positive things to say about the brand.
Having been privy to some of the contractual details of deals that Google has made with other medical centers, I’m betting that they probably got it for ”free”, as in they didn’t directly pay a set fee to the universities. Google most likely provided funding in the form of donations (tax write off), free cloud compute resources and/or cloud storage (write off), and the opportunity for university researchers to co-author high impact publications (everybody wins).
Also, 200k patients is actually kind of small. Granted this dataset is far more granular/robust than what you’d typically find in commercially available healthcare datasets, but to give you some frame of reference, the healthcare datasets I work with contain > 20 million individuals (again, with orders of magnitude fewer features).
Wasn’t that already possible via the rJava library?
That's actually a bit out of date. The nightly builds are now 0.7 (not sure I've seen this mentioned beyond Discourse/Github). It will primarily serve as a depreciation release, i.e. things that would have just broke when going from 0.6 -> 1.0 will instead throw a depreciation warning.
Zero incentive to move to SF. The cost of living has reached a point where I'd end up taking home less disposable income even with a significant increase over my current base salary and signing bonuses get eaten up by moving expenses. Barely escaped the housing bubble crash in south Florida a few years back. No way I'd want to run the risk of experiencing another market adjustment.
Absolutely, though as you mention, removing the ability to use packages and the necessity of writing statistical code that properly accounts for data being spread out across multiple nodes would likely be out of the reach of your everyday/typical R user. An open sourced alternative to Revolution R/Microsoft R Server's out of core processing backend + distributed analtyics packages would be a huge addition to the R language.
I don't believe I praised RStudio at any point in that commment and concern about a very realistic potential hurdle that Rodeo may face as it's codebase grows and matures =/= criticism. I have no issues with the speed of the current implementation (but do have a problem with its lack of features and numerous bugs). It's not a bad IDE, just nowhere near as mature as RStudio.
Realized I never posted the link about R 3.4 that I referenced. https://cdn.ampproject.org/c/s/www.r-bloggers.com/performanc...
Thanks for mentioning it. I almost always forget about Spyder. I'm not sure if that says more about me or Spyder.
Julia's biggest hurdle is the lack of well functioning DataFrames (or the current fork, DataTables). Tons of issues around nullable arrays, etc. have really slowed progress. I do think it's got a ton of upside, but I've found that reimplementing my R or Python scripts in Julia to be too much of a hassle. Costs of reimplemention greatly outweigh the not insignificant gains in speed.
Also check out this article on updates to R 3.4. R tends to be fast enough for most work (I use it regularly on one-off analysis or things that won't ever make it farther than ad-hoc reporting/findings but can't imagine using it in production systems). The listed changes should go a long way towards making R just fast(er) enough for dealing with larger datasets (doesn't help with datasets larger than memory though). For large datasets all the momentum seems to be moving towards Spark (sparklyr is RStudio's SparkR integration. Very much a beta but getting better by the day). On the Python front Dask is awesome for out of memory computation that has no equivalent in R.
Worth mentioning that Rodeo is still unstable and definitely not a feature for feature equivalent for RStudio (I also worry about speed as the size of the project grows as it's based on the same backend as the Atom IDE which has been dogged by speed complaints almost from day one). As for Jupyter Lab, the readme itself says that it's not yet ready for general use. Currently there is no true Python equivalent to RStudio (unfortunately).
The ability to integrate aspects of machine learning into real world production data pipelines and/or the ability to move beyond the creation of the machine learning model itself towards developing an actual app/software suite/product that can be utilized by others (particularly those who have no working knowledge of either programming or machine learning).
I keep looking for another data science spot to open up in Raleigh. Wish I could have taken advantage of the openings when your office first opened up but was terrible timing on my end (really enjoyed interviewing though). Also got to meet Tim not too long ago at PyData Carolinas and he's awesome.