HN user

ctk_brian

117 karma
Posts48
Comments5
View on HN
www.crosstab.io 4y ago

How to use PyTorch LSTMs for time series regression

ctk_brian
3pts0
www.crosstab.io 4y ago

How to use PyTorch LSTMs for time series regression

ctk_brian
2pts0
www.crosstab.io 4y ago

Applications of survival analysis (that aren't clinical research)

ctk_brian
2pts0
www.crosstab.io 4y ago

How to construct survival tables from duration tables

ctk_brian
2pts0
www.crosstab.io 5y ago

How to compute Kaplan-Meier survival curves with SQL

ctk_brian
2pts0
www.crosstab.io 5y ago

How to compute Kaplan-Meier survival curves in SQL

ctk_brian
2pts0
www.crosstab.io 5y ago

A review and how-to guide for Microsoft Form Recognizer

ctk_brian
34pts8
www.axios.com 5y ago

IEA sees gas rebound and issues climate warning

ctk_brian
5pts0
iopscience.iop.org 5y ago

Drivers of greenhouse gas emissions by sector from 1990 to 2018

ctk_brian
1pts0
www.crosstab.io 5y ago

A review and how-to guide for Amazon Textract

ctk_brian
1pts0
www.crosstab.io 5y ago

How to plot survival curves with Plotly and Altair

ctk_brian
2pts0
www.crosstab.io 5y ago

A checklist for professionalizing machine learning models

ctk_brian
3pts0
www.crosstab.io 5y ago

Framing business goals as modeling tasks is the true measure of a data scientist

ctk_brian
1pts0
www.crosstab.io 5y ago

Streamlit review and demo: best of the Python data app tools

ctk_brian
3pts0
www.crosstab.io 5y ago

Google Form Parser, a review and how-to

ctk_brian
1pts0
www.ercot.com 5y ago

ERCOT again asks Texans to reduce electric use

ctk_brian
7pts3
www.noaa.gov 5y ago

May 2021 tied for 6th-warmest May on record for the globe

ctk_brian
2pts0
www.crosstab.io 5y ago

How to build duration tables from event logs with SQL

ctk_brian
2pts0
www.crosstab.io 5y ago

How to convert event logs to duration tables for survival analysis

ctk_brian
1pts0
www.crosstab.io 5y ago

How to build conversion tables from event logs

ctk_brian
1pts0
www.crosstab.io 5y ago

A checklist for professionalizing machine learning models

ctk_brian
2pts0
www.crosstab.io 5y ago

A checklist for professionalizing machine learning models

ctk_brian
1pts0
www.crosstab.io 5y ago

Review: Statistical Rethinking, by Richard McElreath

ctk_brian
1pts0
www.crosstab.io 5y ago

Review: Statistical Rethinking, by Richard McElreath

ctk_brian
2pts0
www.crosstab.io 5y ago

Data before models, but problem formulation first

ctk_brian
1pts0
www.crosstab.io 5y ago

Streamlit app: How to analyze a staged rollout experiment

ctk_brian
2pts0
www.crosstab.io 5y ago

Streamlit review and demo: best of the Python data app tools

ctk_brian
2pts0
share.streamlit.io 5y ago

Show HN: A streamlit app that shows how to analyze a staged rollout

ctk_brian
3pts0
www.crosstab.io 5y ago

How to analyze a staged rollout experiment

ctk_brian
1pts0
www.crosstab.io 5y ago

In Defense of Statistical Modeling

ctk_brian
2pts0

Ah, did I miss that caveat in the documentation somewhere?

What's the use case for that, though? If the documents are highly homogeneous, why would I need a service--let alone an AI service--to extract the data? I could just specify the locations of the fields a priori on a 1040 (for example).

Agreed, but ground truth labeling is a lot of work! The thing is, Form Recognizer has a hard limit of 500 total pages (not documents) in the training set.

I'm skeptical it's possible to achieve good performance with an unsupervised model with only 500 pages, unless those documents are very similar. In which case, why would you need a service like Form Recognizer at all?

From a product perspective, it just makes no sense to me.

I doubt you're wrong. The quote from the ERCOT VP of planning hints that even they are concerned about it:

"We will be conducting a thorough analysis with generation owners to determine why so many units are out of service," said ERCOT Vice President of Grid Planning and Operations Woody Rickerson. "This is unusual for this early in the summer season."