KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Marek Olšák	c8e70e64ac	radeonsi: add flexible shader descriptor management and use it for sampler views It moves all sampler view descriptors to a buffer. It supports partial resource updates and it can also unbind resources (required for FMASK texturing). The buffer contains all sampler view descriptors for one shader stage, represented as an array. On top of that, there are N arrays in the buffer, which are used to emulate context registers as implemented by the previous ASICs (each array is a context). This uses the RCU synchronization approach to avoid read-after-write hazards as discussed in the thread: "radeonsi: add FMASK texture binding slots and resource setup" CP DMA is used to clear the descriptors at context initialization and to copy the descriptors from one context to the next. v2: - use PKT3_DMA_DATA on CIK (I'll test CIK later) - turn the bool CP DMA parameters into self-explanatory flags - add a nice simple API for packet emission to radeon_winsys.h - use 256 contexts, 128 causes texture corruption in openarena	2013-08-17 01:48:25 +02:00
Tom Stellard	764502b481	radeonsi/compute: Let the state tracker do all the flushing It shouldn't be necessary to call radeon_winsys::cs_flush() from radeonsi_launch_grid(), because the state tracker is responsible for flushing the pipeline at the appropriate time. The current behavior is also wrong, because radeonsi_launch_grid() submits packets to the compute ring, but when the state tracker calls pipe->flush() everything is submitted to the graphics ring. This has the potential to create a race condition. The downside of removing this flush is that the compute dispatch packets will be sent to the graphics ring rather than the compute ring. In the future we will need to come up with a way to detect 'compute' command streams and submit them to the appropriate ring. Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2013-08-17 01:48:25 +02:00
Kenneth Graunke	e29931aa74	i965: Dump more information about batch buffer usage. Previously, INTEL_DEBUG=bat would dump messages like: intel_mipmap_tree.c:1643: Batchbuffer flush with 456b used This only reported the space used for command packets, and didn't report any information on the space used for indirect state. Now it dumps: intel_context.c:366: Batchbuffer flush with 6128b (pkt) + 4288b (state) = 10416b (31.8%) This conveniently shows the breakdown of space used for packets vs. state, as well as the percentage of batchbuffer space. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 15:54:24 -07:00
Kenneth Graunke	2a9492f321	i965: Add Gen7 depth stall flushes before disabling depth in BLORP. We emit these before configuring depth in the normal path, or actually using the depth buffer in BLORP - we just failed to emit them when disabling depth altogether. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Chad Versace <chad.versace@linux.intel.com> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2013-08-16 15:03:55 -07:00
Kenneth Graunke	8fba8d4ee7	i965: Add Gen6 depth stall flushes before disabling depth in BLORP. We emit these before configuring depth in the normal path, or actually using the depth buffer in BLORP - we just failed to emit them when disabling depth altogether. On Sandybridge, this also requires the post_sync_nonzero flush. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Chad Versace <chad.versace@linux.intel.com> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2013-08-16 15:03:38 -07:00
Matt Turner	9c48ae751a	i965: Don't copy propagate bitcasts with source modifiers. Previously, copy propagation would cause bitcast_f2u(abs(float)) to be performed in a single step, but the application of source modifiers (abs, neg) happens after type conversion, leading to incorrect results. That is, for bitcast_f2u(abs(float)) we would in fact generate code to do abs(bitcast_f2u(float)). For example, whereas bitcast_f2u(abs(float)) might result in a register argument such as (abs)g2.2<0,1,0>UD v2: Set interfered = true and break in register_coalesce instead of returning false. Reviewed-by: Paul Berry <stereoytpe441@gmail.com>	2013-08-16 13:11:07 -07:00
Matt Turner	0ae9ca12a8	i965: Emit MOVs for neg/abs. Necessary to avoid combining a bitcast and a modifier into a single operation. Otherwise if safe, the MOV should be removed by copy-propagation or register coalescing. With this and the next patch, there are only four changes in shader-db: all a single extra instruction. The code does something like mov a.w, -b.x and copy propagation doesn't work because it only handles no-op swizzles. Seems acceptable, given the known limitation of our copy propagation. Reviewed-by: Ian Romanick <ian.d.romanick@intel.com> Reviewed-by: Paul Berry <stereoytpe441@gmail.com>	2013-08-16 13:11:07 -07:00
Anuj Phogat	079bdba05f	i965/blorp: Add support for single sample scaled blit with bilinear filter Currently single sample scaled blits with GL_LINEAR filter falls back to meta path. Patch removes this limitation in BLORP engine and implements single sample scaled blit with bilinear filter. No piglit, gles3 regressions are observed with this patch on Ivybridge. V2: Use "sample" message to utilize the linear filtering functionality built in to hardware. V3: Define a bool variable (bilinear_filter) to handle the conditions for GL_LINEAR blits. Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 09:46:15 -07:00
Anuj Phogat	aff371b634	i965/blorp: Define a function to clamp texture coordinates New function clamp_tex_coords() clamps the texture coordinates to texture boundaries. This function will also be utilized later for the BLORP implementation of single-sample scaled blit with bilinear filter. Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 09:46:15 -07:00
Anuj Phogat	6066fb1721	i965/blorp: Use more appropriate variable names When we talk about both multi-sample and single-sample scaled blits, rect_grid_{x1, y1} are more appropriate variable names as compared to sample_grid_{x1, y1}. There are no functional changes in this patch. It just prepares for the BLORP implementation of single-sample scaled blit with bilinear filter. Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 09:46:15 -07:00
Anuj Phogat	d944a6144f	meta: Fix blitting a framebuffer with renderbuffer attachment This patch fixes a case of framebuffer blitting with renderbuffer as color attachment and GL_LINEAR filter. Meta implementation of glBlitFrambuffer() converts source color buffer to a texture and uses it to do the scaled blitting in to destination buffer. Using the exact source rectangle to create the texture does incorrect linear filtering along the edges. This patch makes the changes to extend the texture edges by one pixel in x, y directions. This ensures correct linear filtering. It fixes failing piglit fbo-attachments-blit-scaled-linear test. Signed-off-by: Anuj Phogat <anuj.phogat@gmail.com> CC: "9.2" <mesa-stable@lists.freedesktop.org> CC: "9.1" <mesa-stable@lists.freedesktop.org> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 09:46:15 -07:00
Ilia Mirkin	a2061eea0f	nv50: add vp3/vp4 support for mpeg2/vc1 h264/mpeg4 remain disabled for pre-nvc0, there's some minor bug/difference which causes the decoding to hang after some frames. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-16 09:48:47 +02:00
Ilia Mirkin	b3f6f127f2	nv50: separate video logic from noalloc The upcoming vp3 logic will want the video layout, but allocated by the miptree. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-16 09:48:26 +02:00
Ilia Mirkin	c1a6f59b20	nv30: remove no-longer-used formats from table Commit `14ee790df7` removed the formats from the vtxfmt_table but forgot to also update the info_table. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "9.2 and 9.1" <mesa-stable@lists.freedesktop.org>	2013-08-16 09:48:09 +02:00
Fredrik Höglund	0e7a61a29f	mesa: Update the BGRA vertex array error handling The error code was changed from INVALID_VALUE to INVALID_OPERATION in OpenGL 3.3. We should also generate an error when size is BGRA and normalized is FALSE. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2013-08-15 21:38:13 -07:00
Kenneth Graunke	90129da82c	i965/fs: Fix Sandybridge regressions from SEL optimization. Sandybridge is the only platform that supports an IF instruction with an embedded comparison. In this case, we need to emit a CMP to go along with the SEL. Fixes regressions in Piglit's glsl-fs-atan-3, fs-unpackHalf2x16, fs-faceforward-float-float-float, isinf-and-isnan fs_basic, and isinf-and-isnan fs_fbo. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=68086 Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Matt Turner <mattst88@gmail.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com> Tested-by: lu hua <huax.lu@intel.com>	2013-08-15 15:33:00 -07:00
Kenneth Graunke	c189840b21	i965: Force X-tiling for 128 bpp formats on Sandybridge. 128 bpp formats are not allowed to be Y-tiled on any architectures except Gen7. +11 Piglits on Sandybridge (mostly regression fixes since the switch to Y-tiling). Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=63867 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=64261 Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Chad Versace <chad.versace@linux.intel.com> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com> Cc: "9.2" <mesa-stable@lists.freedesktop.org>	2013-08-15 15:18:48 -07:00
Ian Romanick	41eef83cc0	mesa/vbo: Fix handling of attribute 0 in non-compatibilty contexts It is only in OpenGL compatibility-style contexts where generic attribute 0 and GL_VERTEX_ARRAY have a bizzare, aliasing relationship. Moreover, it is only in OpenGL compatibility-style contexts and OpenGL ES 1.x where one of these attributes provokes the vertex. In all other APIs each implicit call to glArrayElement provokes a vertex regardless of which attributes are enabled. Signed-off-by: Ian Romanick <ian.d.romanick@intel.com> Reviewed-by: Robert Bragg <robert@sixbynine.org> Cc: "9.0 9.1 9.2" <mesa-stable@lists.freedesktop.org> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=55503 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=66292 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=67548	2013-08-15 14:59:37 -07:00
Zack Rusin	7115bc3940	draw: handle nan clipdistance If clipdistance for one of the vertices is nan (or inf) then the entire primitive should be discarded. Signed-off-by: Zack Rusin <zackr@vmware.com> Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2013-08-15 16:26:32 -04:00
Vinson Lee	035bf21983	i915,i965: Fix memory leak in try_pbo_upload (v2) Fixes "Resource leak" defect reported by Coverity. Tested on Haswell, no Piglit regressions. v2: Apply to i965, not just i915. (chadv) CC: "9.2, 9.1" <mesa-stable@lists.freedesktop.org> Signed-off-by: Vinson Lee <vlee@freedesktop.org> Reviewed-by: Chad Versace <chad.versace@linux.intel.com>	2013-08-15 10:37:22 -07:00
Roland Scheidegger	6ca18e06ae	gallivm: revert accidentally commited hunk That magic wasn't meant to be commited, need to work on some proper fix.	2013-08-15 19:26:39 +02:00
Roland Scheidegger	5626a84a00	gallivm: do per-sample depth comparison instead of doing it post-filter Doing the comparisons pre-filter is highly recommended by OpenGL (and d3d9) and definitely required by d3d10. This actually doesn't do it pre-filter but more "in-filter" as otherwise need to push the comparisons even further down into fetch code and this also trivially allows using a somewhat cheaper lerp. Doing it pre-filter would actually have some performance advantage for UNORM formats (because the comparisons should be done in texture format, we'd only need to convert the shadow ref coord to texture format once, but in turn would save converting the per-sample texture values to floats) but this gets a bit messy as this has implications for border color handling as well (which needs to be done prior to depth comparisons, hence would also need to convert border color to texture format too or use some other tricks like doing separate border color / shadow ref comparison and simply using that result directly when doing border replacement). Should make no difference for nearest filtering, and performance for linear filtering should be mostly the same too (essentially have one more comparison instruction per sample, and replace the sub/mul/add lerp with a sub/and/and/add special "lerp" which all in all shouldn't be much of a difference). v2: get rid of old code completely Reviewed-by: Zack Rusin <zackr@vmware.com>	2013-08-15 18:42:20 +02:00
Michel Dänzer	3b2f3f90ac	radeonsi: Pixel shaders pre-load one more SGPR Acked-by: Marek Olšák <maraeo@gmail.com>	2013-08-15 17:55:00 +02:00
Michel Dänzer	f0753a3cd4	radeonsi: TGSI_SEMANTIC_CLIPVERTEX doesn't use any parameters	2013-08-15 17:54:40 +02:00
Michel Dänzer	2f98dc223f	radeonsi: Don't export unused clip distance vectors from vertex shader E.g. the Source engine seems to always write to gl_ClipVertex, but normally doesn't enable any GL_CLIP_DISTANCEn states. This change removes some irrelevant parts from the generated vertex shader code in such cases. Reviewed-by: Tom Stellard <thomas.stellard@amd.com>	2013-08-15 17:53:50 +02:00
Michel Dänzer	b00269aa58	radeonsi: Don't leave gaps between position exports from vertex shader If the vertex shader exports clip distances but not point size, use position exports 1/2 instead of 2/3 for the clip distances. Fixes geometry corruption in that case. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=66974 Cc: mesa-stable@lists.freedesktop.org Reviewed-by: Tom Stellard <thomas.stellard@amd.com>	2013-08-15 17:42:26 +02:00
Roland Scheidegger	abdd32dcd5	llvmpipe: fix stencil bug if we have both stencil and depth tests This is a very well hidden bug found by accident (only the fixed glean tstencil2 test so far seems to hit it). We must use new mask with combined s_pass values and orig_mask values for zpass/zfail stencil ops, otherwise both the sfail op and one of zpass/zfail op are applied (probably not hit in most tests because some of the ops tend to be KEEP usually). Note: this is a candidate for the 9.2 branch. Reviewed-by: Zack Rusin <zackr@vmware.com>	2013-08-15 17:30:07 +02:00
Roland Scheidegger	7ae9cc71f0	st/mesa: use new float comparison opcodes if native integers are supported Should get rid of some float-to-int conversions (with negation). No piglit regressions (with llvmpipe). v2: fix bogus formatting spotted by Brian. Reviewed-by: Brian Paul <brianp@vmware.com>	2013-08-15 17:30:07 +02:00
Ilia Mirkin	4ea191fb2d	nvc0: move video param and format support functions to nouveau Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:48 +02:00
Ilia Mirkin	9255019a53	nvc0: move firmware loading functions to nouveau Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:48 +02:00
Ilia Mirkin	9d8c076803	nvc0: move some of the simpler decoder functions into nouveau Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:48 +02:00
Ilia Mirkin	73f4499a02	nvc0: move vp param filling logic into nouveau Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:48 +02:00
Ilia Mirkin	e1cd987bb6	nvc0: move bsp param-filling logic into nouveau Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:48 +02:00
Ilia Mirkin	d6a82a7747	nvc0: move nvc0_decoder into nouveau, rename to nouveau_vp3_decoder Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:47 +02:00
Ilia Mirkin	86e5c3c97b	nvc0: standardize on using #if for NVC0_DEBUG_FENCE Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:47 +02:00
Ilia Mirkin	b57875bbb3	nvc0: refactor video buffer management logic into nouveau_vp3 Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:19:47 +02:00
Ilia Mirkin	940f7cec77	nv50: allow forcing PMPEG use, for ease of testing This also allows people who don't want to install the binary blobs required for VP2 to still get MPEG decoding. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:15:23 +02:00
Ilia Mirkin	ee3ca3614e	nv30: hook up PMPEG support via nouveau_video, enables XvMC to work Force the format to be the reasonable format that doesn't require an inverse z-scan. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:15:12 +02:00
Ilia Mirkin	6010c683d0	nouveau: set buffer format of video buffer Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:15:04 +02:00
Ilia Mirkin	8975f83402	nouveau: fix number of surfaces in video buffer, use defines Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2013-08-15 15:15:02 +02:00
Ilia Mirkin	14ee790df7	nv30: U8_USCALED only works for size 4 See https://bugs.freedesktop.org/show_bug.cgi?id=61635 for a sample program. Changing it to use a vec4 makes it work. Remove the unsupported formats. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "9.2 and 9.1" <mesa-stable@lists.freedesktop.org>	2013-08-15 15:14:25 +02:00
Chris Forbes	4f739646b0	i965: allow 8 user clip planes on CTG+ There's no need to use a clip flag for NEGW on these gens, so no reason we can't just enable 8 planes. V2: - Bump (and document!) MAX_VERTS in the clip code. - Fix clip flag masks in the clip unit state and in the shader prolog - Move this to the end of the series for less breakage. Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:56 +12:00
Chris Forbes	ee0b8e0f06	i965: get rid of clip plane compaction Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:56 +12:00
Chris Forbes	cf52f6435e	i965/clip: Support clip distances for line clipping This does the same thing as we do for triangle clipping -- select the appropriate source (either dot(hpos,fixed plane) or a clipdistance slot). Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:56 +12:00
Chris Forbes	2a8a85e1ad	i965/clip: remove spurious clipvertex param Nothing in the clipper uses gl_ClipVertex any more, so we don't care where it is. V2: Don't bother fishing out the clipvertex offset either. Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:56 +12:00
Chris Forbes	45540921ec	i965/clip: Use clip distances for all user clipping V2: Adjust explanation of load_clip_distance() Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:55 +12:00
Chris Forbes	bf9ede92c2	i956/clip: push dp4 into load_clip_distance Soon the dp4 is only going to be used for fixed clip planes. V2: Remove old inaccurate comment about the behavior of this function; add a better explanation above. Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:55 +12:00
Chris Forbes	265336e75a	i965/clip: Track offset into the vertex for clipdistance Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:55 +12:00
Chris Forbes	3b738f5f85	i965/Gen4-5: Set clip flags from clip distances V2: - Use the new VS_OPCODE_UNPACK_FLAGS_SIMD4X2 to correctly split the flags for the two vertices being processed together. - Don't apply bogus masking of clip flags. The set of plane enables aren't included in the shader key, and we wouldn't want the recompiles anyway. V3: - Tidy up spurious instructions, name temps properly. Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> [V2] Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:55 +12:00
Chris Forbes	a9be50f776	i965: add new VS_OPCODE_UNPACK_FLAGS_SIMD4X2 Splits the bottom 8 bits of f0.0 for further wrangling in a SIMD4x2 program. The 4 bits corresponding to the channels in each program flow are copied to the LSBs of dst.x visible to each flow. This is useful for working with clipping flags in the VS. V3: - Fixup immediate types - Teach scheduler about the hidden dep on flags Signed-off-by: Chris Forbes <chrisf@ijw.co.nz> V2: Reviewed-by: Paul Berry <stereotype441@gmail.com>	2013-08-16 07:24:38 +12:00

1 2 3 4 5 ...

58070 Commits All Branches Search

58070 Commits

All Branches