Hey folks - I'm a Java Champion and a maintainer on LangChain4j. I have been working in the eval space for close to 2 years now. I came across arize.com/docs/…/openinference-java and wanted to chat with someone about how the openinference sdk differs from opentelemetry and the OpenTelemetry semantic conventions. I was also wondering about direct integration with Arize Phoenix via java. I contributed a java API for langfuse (github.com/langfuse/langfuse-java/pull/36) as well as I own a quarkus extension for langfuse (docs.quarkiverse.io/quarkus-langfuse/dev) and was wondering if anyone from Arize Phoenix would be interested in something similar for Arize?
Hey Eric, good to have you here. Would love your contributions to github.com/Arize-ai/openinference the conventions are a bit more AI and agent forward, though Phoenix and AX do support GenAI so the lines are getting blurry at this point. Take a look and let us know what you think! could use help on the Java side
thanks Mikyo - i've been reading through openinference and to be honest I've been having a hard time seeing the line between what its providing vs what is in OpenTelemetry
langchain4j itself is instrumented with opentelemetry, and otel itself has genai semantic conventions
i've also just proposed github.com/Arize-ai/client_java/issues/16
Thanks for taking the time. Some context on where I'm coming from: I've standardized on OpenTelemetry. My stack is Java/Quarkus + LangChain4j, and Quarkus LangChain4j already emits `gen_ai.*` spans natively. Everything flows through an OTel Collector, and I want the backend to be a choice, not a re-instrumentation decision. So my core question is: **what would I lose if I only ever emitted OTel GenAI conventions and pointed them at Phoenix/AX?** A concrete list would be the most useful answer. More specific questions: 1. Coexistence contract
If a span carries both `gen_ai.*` and OpenInference attributes, which is authoritative, in Phoenix and in AX? (The AX docs say OpenInference wins. Is that the same in Phoenix OSS?)
Is dual emission a supported, documented pattern, or a transitional hack?
Which OpenInference fields does Phoenix actually *require* for full rendering, evals, and datasets, versus optional enrichment?
Which concepts do you consider intentionally non-mappable to `gen_ai.*`? (github.com/Arize-ai/openinference/issues/2130 asks for this. Is there a target date?)
2. Phoenix `gen_ai.*` ingestion fidelity
The 15.10 conversion table maps `gen_ai.usage.*` to prompt/completion counts only. Are cache read/write, reasoning, and per-modality tokens mapped?
It keys provider off `gen_ai.system`. Is `gen_ai.provider.name` handled? (`gen_ai.system` is deprecated upstream.)
Does Phoenix infer span kind from `gen_ai.operation.name` like AX does, including `retrieval`, `execute_tool`, `invoke_agent`, `invoke_workflow`, and `plan`? The table shows operation name landing in `llm.invocation_parameters`.
Are `gen_ai.evaluation.result` events ingested as annotations?
How do you track spec churn? For example, the `cache_creation` vs `cache_write` naming drift in the current GenAI docs.
3. Emitting `gen_ai.*` from OpenInference (#2817, #2983)
What's the status and timeline for the bidirectional layer / "emit GenAI format" switch?
Will it cover Java? If I use `openinference-instrumentation-langchain4j`, can I get conformant `gen_ai.*` out of it?
Would you consider adding OpenInference Java instrumentors to the OTel GenAI conformance dashboard (Weaver live-check)? Today only Python and JS appear.
4. Java specifics
Quarkus LangChain4j already creates spans for AI service methods, chat model calls, tools, and guardrails. Does the OpenInference LangChain4j instrumentor create *additional* spans alongside those, or can it enrich the existing ones? What's the recommended setup to avoid duplicates?
Which Maven coordinates are canonical, `com.arize:*` (0.1.x) or `io.openinference:*` (1.0.0)? The docs show both.
The LangChain4j instrumentor is 0.1.x. How does that square with "stable conventions"? Is the stability guarantee about the spec, the SDKs, or both?
5. Gaps where you're ahead of OTel, and upstreaming
RAG: OTel retrieval docs carry only id + score, with no reranker operation. Are you proposing `document.content`/`metadata` and a reranker operation upstream, or keeping those OpenInference-only?
Are `CHAIN`, `GUARDRAIL`, `PROMPT`, and `EVALUATOR` span kinds, the `llm.cost.*` tree, and the reasoning-replay fields (signatures, encrypted content) being proposed to the GenAI SIG? Is anyone from Arize active there?
If OTel GenAI goes stable with different shapes for these, what's the migration story for existing OpenInference data and dashboards?
6. Privacy defaults
OpenInference captures content unless `OPENINFERENCE_HIDE_*` is set; OTel GenAI makes content opt-in. Would you consider an opt-in default, or a mode that follows OTel's content-capture guidance, including the external-storage pattern?
7. Governance
How are spec changes proposed and decided, and is there a versioned changelog for the spec itself, separate from the SDKs?
Beyond Phoenix and AX, which backends consume OpenInference natively? Dynatrace, for example, documents converting OpenInference *to* `gen_ai.*`.
Happy to help test any of this against Quarkus LangChain4j.
FYI I'm also the person who authored all of the dev.langchain4j.observability.api stuff in LangChain4j
Hey Eric, just acking here. I'll respond with answers when I free up
no worries! I'm also brand new to arize, but its something thats gotten my interest working in the open source eval space as a LangChain4j & Quarkus maintainer, and more broadly as someone trying to promote Java in the AI space. I work with lots of large enterprises who standardize on Java and aren't going to switch to python for their applications, and I'm finding the overall space lacking in industry from the Java integration side of things.
and didn't mean to overwhelm 🙂
