I have been working on bootstrapping AI evals before product launch and came up with something like this. Using QA to inform "problematic on purpose" synthetic data that can validate LLM-as-a-Judge. Would love to chat with people who have experience with this :) #mlsky #buildinpublic
2 likes 0 replies
?