How are you operating AI infrastructure in production?

6 points by kkgupta ↗ HN
There are many open-source projects across inference, orchestration, observability, vector search, data pipelines, evaluation, and model management. Most are relatively easy to test, but production operation is a different problem.

For those running open-source AI infrastructure in production:

- What are you running, for what workload, and would you recommend?

- Do you operate yourself versus consume as a managed service?

- Have you replaced or abandoned any tools because they were too difficult or expensive to operate?

- What problems only appeared after moving beyond the prototype stage?

- Anything that you would do differently if rebuilding the stack today?

Thanks

0 comments

[ 0.35 ms ] story [ 8.3 ms ] thread

No comments yet.