HN user

kprybol

107 karma
Posts0
Comments35
View on HN
No posts found.

Counterpoint to that is when a recruiter initiates contact they are often very coy until they have you on the phone, frequently not even sending a JD until after that first conversation (where this line of questioning usually arises). I get wanting to provide an engaging response but hard to do so until the recruiter has provided sufficient information.

Certilytics | Hiring Data Engineers, Product Managers, Analysts, Developers, and more | Remote: US Only | Full-time

Certilytics provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights. Our team represents a dynamic infusion of multidiscipline, which includes actuarial, data, and behavioral scientists, IT engineers, software developers, nurse clinicians, and experts in public health and the health insurance industry. Certilytics has extensive experience working with a diverse set of customers, including large self-insured employers, health plans, pharmacy benefit managers, government programs, care management companies, and health systems. These relationships with various data providers and customers allow for rapid data ingestion, validation, and enrichment, as well as streamlined delivery of analytic dashboards, outputs, and visualizations to our customers. Our unique approach allows for developing the most accurate financial, clinical and behavioral models in the industry.

Why Certilytics! Access to one of the most extensive clinical datasets in the industry that includes medical claims, pharmacy claims, and laboratory data. Impactful work. We're big enough to have the freedom to take on interesting projects but small enough that your work is always important and highly visible within the organization. Remote friendly. The Certilytics team is distributed throughout the US and has regular in-person working sessions for all of those little things that are hard to accomplish over Teams.

See our open jobs: https://jobs.silkroad.com/CarewiseHealth/Certilytics

Certilytics | Hiring Data Scientists and Data Engineers | Remote: US Only | Full-time

Certilytics provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights. Our team represents a dynamic infusion of multidiscipline, which includes actuarial, data, and behavioral scientists, IT engineers, software developers, nurse clinicians, and experts in public health and the health insurance industry. Certilytics has extensive experience working with a diverse set of customers, including large self-insured employers, health plans, pharmacy benefit managers, government programs, care management companies, and health systems. These relationships with various data providers and customers allow for rapid data ingestion, validation, and enrichment, as well as streamlined delivery of analytic dashboards, outputs, and visualizations to our customers. Our unique approach allows for developing the most accurate financial, clinical and behavioral models in the industry.

Why Certilytics!

Access to one of the most extensive clinical datasets in the industry that includes medical claims, pharmacy claims, and laboratory data.

Impactful work. We're big enough to have the freedom to take on interesting projects but small enough that your work is always important and highly visible within the organization.

Remote friendly. The Certilytics data science team is distributed throughout the US and has regular in-person working sessions for all of those little things that are hard to accomplish over Teams.

Tech Stack: Python, TensorFlow, Kubernetes, AWS, Spark, Scala

See our open jobs: https://www.certilytics.com/contact-us/#careers

Certilytics | Louisville, KY | REMOTE (US Only) | Machine Learning Research Engineer

Do you enjoy reading the latest machine learning research on Arxiv? Do you challenge yourself to reverse engineer interesting papers? Do you seek to apply existing algorithms to new domains and develop creative and novel solutions to difficult problems? If so, come join our team at Certilytics!

Certilytics, Inc. provides sophisticated predictive analytics solutions to major healthcare organizations by integrating financial, clinical, and behavioral insights.

As a machine learning research engineer, you'll be responsible for designing and running experiments to bring the latest deep learning advances from the literature to our products. As part of the data science team, you will be responsible for building models for clinical and financial risk prediction, performing original research, and contributing to a proprietary machine learning library. The ideal candidate will have a strong background in natural language processing and familiarity with the inner workings of RNN’s and transformer networks (Join a flexible, energetic team in bringing the best of deep learning to healthcare.

Apply here: https://jobs.silkroad.com/CarewiseHealth/Certilytics/jobs/20...

From my experiences (currently work with several Fortune 100 health insurers/benefits managers, and have previously worked for another large insurer, a major academic medical center, and a large pharma company), healthcare organizations tend to be rather cloud adverse (most of our contracts very explicitly forbid us from using any form of 3rd party cloud computing). So while I agree that much of the heavy lifting will shift to the cloud (or already has), I expect health analytics will continue to favor on-premises solutions (GPU’s still tend to be pretty rare compared to CPU based clusters but are slowly becoming more common).

That's fair. Google's recent paper on predicting patient deaths is another good example of this (logistic regression + good feature engineering performed just as well as their deep learning models, and the logistic regression has the added benefit of being significantly more interpretable and as a result, actionable).

It'll be interesting to see when specialized ML focused silicon will become readily available. Right now I find ML libraries that are able to run on blended architectures (any combination of CPU and GPU's) much more exciting/impactful than TPU's. The ability to deploy on just about any cluster a customer may have available is huge.

I've always viewed DeepMind as more of a skunk works program and less as a profit driven enterprise. DeepMind exists primarily to push the limits of what can be done when you put group of leading researchers together in a room, provide them with nearly limitless resources, and simply tell them to "go". I expect some of that effort to eventually trickle down into Google's consumer products (maybe a healthcare focused version of AutoML https://cloud.google.com/automl/). Google has already done a lot of work on the HIPPA side of things (https://cloud.google.com/security/compliance/hipaa/)

Being rules based isn't necessarily a bad thing or disingenuous. I develop healthcare AI products (ML/DL researcher) and we actually aim to be able to translate our models into a rules based engine (find a strong signal, interpret/understand model well enough to translate/embed into a rules engine, look for a new signal in our models, rinse + repeat). We end up deploying a mix of rules based and true ML based models into production but it may not be immediately obvious to the end user which type of model they are using.

Healthcare data can't be shared the way the browser histories, cell phone location data, etc. can. It's a completely different set of a rules that people have to play by (HIPPA for example). I build machine learning systems using healthcare claims and EHR data and without the direct cooperation of several large insurance companies (and access to their data) we'd be dead in the water. Even with access to their data there are incredibly strict limits to what we can and can't do with it. You can't just go out and collect healthcare data the way you can many other types of data.

Towards Scala 3 8 years ago

It depends on the use case. Our work primarily revolves around extending Spark with custom pipelines, models, ensembles, etc. to be deployed into our production systems (petabyte scale). Scala was really the only way to go for us.

Towards Scala 3 8 years ago

Scala is used pretty heavily in the big data world, particularly if you are working with Spark.

Not sure how easy it is to find outside of boutique pets stores (just happen to live near a fantastic one that delivers for free) but I highly recommend the Orijen brand. Have tried numerous other “premium” brands with my dogs over the years but no matter which brand I was using, I always managed to run into at least one vet than had less than stellar things to say about brand x, y, or z. Maybe it’s just coincidence but in the 7+ years I’ve been feeding my dogs Orijen I’ve yet to encounter a vet that anything other than positive things to say about the brand.

Having been privy to some of the contractual details of deals that Google has made with other medical centers, I’m betting that they probably got it for ”free”, as in they didn’t directly pay a set fee to the universities. Google most likely provided funding in the form of donations (tax write off), free cloud compute resources and/or cloud storage (write off), and the opportunity for university researchers to co-author high impact publications (everybody wins).

Also, 200k patients is actually kind of small. Granted this dataset is far more granular/robust than what you’d typically find in commercially available healthcare datasets, but to give you some frame of reference, the healthcare datasets I work with contain > 20 million individuals (again, with orders of magnitude fewer features).

Julia 0.6 is out 9 years ago

That's actually a bit out of date. The nightly builds are now 0.7 (not sure I've seen this mentioned beyond Discourse/Github). It will primarily serve as a depreciation release, i.e. things that would have just broke when going from 0.6 -> 1.0 will instead throw a depreciation warning.

Zero incentive to move to SF. The cost of living has reached a point where I'd end up taking home less disposable income even with a significant increase over my current base salary and signing bonuses get eaten up by moving expenses. Barely escaped the housing bubble crash in south Florida a few years back. No way I'd want to run the risk of experiencing another market adjustment.

Absolutely, though as you mention, removing the ability to use packages and the necessity of writing statistical code that properly accounts for data being spread out across multiple nodes would likely be out of the reach of your everyday/typical R user. An open sourced alternative to Revolution R/Microsoft R Server's out of core processing backend + distributed analtyics packages would be a huge addition to the R language.

I don't believe I praised RStudio at any point in that commment and concern about a very realistic potential hurdle that Rodeo may face as it's codebase grows and matures =/= criticism. I have no issues with the speed of the current implementation (but do have a problem with its lack of features and numerous bugs). It's not a bad IDE, just nowhere near as mature as RStudio.

Julia's biggest hurdle is the lack of well functioning DataFrames (or the current fork, DataTables). Tons of issues around nullable arrays, etc. have really slowed progress. I do think it's got a ton of upside, but I've found that reimplementing my R or Python scripts in Julia to be too much of a hassle. Costs of reimplemention greatly outweigh the not insignificant gains in speed.

Also check out this article on updates to R 3.4. R tends to be fast enough for most work (I use it regularly on one-off analysis or things that won't ever make it farther than ad-hoc reporting/findings but can't imagine using it in production systems). The listed changes should go a long way towards making R just fast(er) enough for dealing with larger datasets (doesn't help with datasets larger than memory though). For large datasets all the momentum seems to be moving towards Spark (sparklyr is RStudio's SparkR integration. Very much a beta but getting better by the day). On the Python front Dask is awesome for out of memory computation that has no equivalent in R.

Worth mentioning that Rodeo is still unstable and definitely not a feature for feature equivalent for RStudio (I also worry about speed as the size of the project grows as it's based on the same backend as the Atom IDE which has been dogged by speed complaints almost from day one). As for Jupyter Lab, the readme itself says that it's not yet ready for general use. Currently there is no true Python equivalent to RStudio (unfortunately).

The ability to integrate aspects of machine learning into real world production data pipelines and/or the ability to move beyond the creation of the machine learning model itself towards developing an actual app/software suite/product that can be utilized by others (particularly those who have no working knowledge of either programming or machine learning).

I keep looking for another data science spot to open up in Raleigh. Wish I could have taken advantage of the openings when your office first opened up but was terrible timing on my end (really enjoyed interviewing though). Also got to meet Tim not too long ago at PyData Carolinas and he's awesome.